2026-10-11 17:11 UTC

Independent replication and audit will determine whether Ox Alpha can reproducibly resolve roughly 96% of SWE-bench Verified-Mini under the official mini-swe-agent scaffold without leakage or evaluation errors.

state: expiredheat: lowuncertainty: highconvergesscott: highcoding-agents swe-bench benchmark-validationOx AlphaSWE-bench

What is this?

Ox Alpha is reported in the case evidence as resolving 48 of 50 tasks (96%) on SWE-bench Verified Mini using the official mini-swe-agent scaffold. The supplied sources establish that Verified Mini has a HAL leaderboard for reproduced results and that mini-swe-agent configurations and versions can materially affect comparability; they also flag contamination and benchmark-quality risks in SWE-bench Verified. However, none of the supplied snippets specifically documents Ox Alpha’s run or an independent replication, so the 96% result and the absence of leakage or evaluation errors remain unverified here.

Why it matters to Scott

The claimed 96% result, coupled with the author’s call for replication and audit, converges directly with Scott’s positions that coding capability must be attributed to a disclosed model-plus-harness unit and validated through trace-backed, independently different checks. If reproduced without contamination, it would also provide consequential evidence that SWE-bench Verified Mini is nearing saturation; if it fails, the failure would expose precisely the harness, leakage, or evaluator risks his frameworks anticipate.
ip:concept.model-plus-harness-benchmark-unitip:concept.future-leakage-ruleip:concept.mechanically-different-verifiersdev:concept.trace-backed-agent-comparisonradar:concept.agent-benchmarksradar:concept.benchmark-integrityradar:concept.agent-evaluationradar:ai-benchmark-saturation-distortion
queries asked of Scott's wikis
  • coding-agent benchmark reproducibility and harness control
  • SWE-bench contamination leakage and evaluation validity
  • mini-swe-agent scaffold effects on model performance
  • independent replication standards for agent benchmarks
  • coding-agent capability saturation and benchmark replacement
  • benchmark scores versus real-world software engineering performance

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (9) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐I benchmarked Ox Alpha on SWE-bench Verified Mini (50 tasks): 96% resolved. Now I’m skeptical of myself.
LocalLLaMA
No_Tip99173517
🟧 hnOx Alpha Is Performing SOTA and as Well as GPT-5.6 Sol in Multi-Agent Arenasensho11
🟠 redditListen guys is gemini cooking wth ???? Ox alpha is it 3.5 pro ???
singularity
Independent-Wind44621028
🟠 redditDeepmind Researcher Strongly Hints Ox Alpha Is The Next Gemini Pro Model
singularity
Neurogence27470
🟠 redditFound out the model behind Ox Alpha. It's unreleased z.ai's GLM model
singularity
py_blu7425
🟠 redditstealth/ox-alpha
LocalLLaMA
danigoncalves015
🟧 hnOx-Alpha Is GLMjitbit7754
🟠 redditOx-alpha: pelican on bicycle benchmark
singularity
TensorFlar21339
🟠 redditOx Alpha erroring out for me.
singularity
Aggravating-Push-20729

Interpretation history

Decision trace