Ox Alpha is reported in the case evidence as resolving 48 of 50 tasks (96%) on SWE-bench Verified Mini using the official mini-swe-agent scaffold. The supplied sources establish that Verified Mini has a HAL leaderboard for reproduced results and that mini-swe-agent configurations and versions can materially affect comparability; they also flag contamination and benchmark-quality risks in SWE-bench Verified. However, none of the supplied snippets specifically documents Ox Alpha’s run or an independent replication, so the 96% result and the absence of leakage or evaluation errors remain unverified here.
2026-08-25T08:29:57Z
Another discussion refresh adds no replication, inspectable artifacts, leakage analysis, or evaluator audit after several days of repetitive attribution speculation. The initial attention window has closed without validation, so the case should expire and only reopen on auditable benchmark evidence.
2026-08-25T07:31:48Z
The refreshed discussion only revises attribution anecdotes and includes a comment-count regression consistent with source refresh noise; it adds no replication, artifacts, traces, leakage analysis, or evaluator audit. The 96% result remains an isolated benchmark claim and should stay dormant pending auditable evidence.
2026-08-25T05:27:44Z
The refreshed HN comment and engagement add no replication, inspectable run artifacts, leakage analysis, or evaluator audit. The 96% result remains an isolated claim; discussion churn should not revive the case without auditable benchmark evidence.
2026-08-25T03:31:40Z
The latest comment and engagement refresh adds only repeated model-attribution discussion, with no independent replication, inspectable run artifacts, traces, leakage analysis, or evaluator audit. The 96% result remains an isolated benchmark claim and should stay dormant pending auditable evidence.
2026-08-25T01:23:44Z
The comment refresh and engagement movement add no independent replication, inspectable artifacts, traces, leakage analysis, or evaluator audit; the apparent comment-count regression is a source refresh artifact, not evidence. The 96% result remains an isolated claim and should stay dormant pending auditable validation.
2026-08-24T23:32:55Z
The refreshed comments and engagement remain repetitive provenance, capability, and reliability anecdotes; they add no matching replication, inspectable run artifacts, traces, leakage analysis, or evaluator audit. The 96% result therefore remains an isolated benchmark claim with no changed meaning.
2026-08-24T22:33:20Z
The isolated error report adds no documented outage pattern and does not validate or challenge the claimed SWE-bench run. The case remains a single-source benchmark anomaly and should stay dormant until a comparable replication, inspectable artifacts, or evaluator audit appears.
2026-08-24T22:23:08Z
evidence attached: reddit.post.1vxgpp0 — An independent user reports Ox Alpha errors, adding weak operational evidence that uptime and reliability may qualify its headline benchmark result.
2026-08-24T21:37:27Z
The pelican-on-bicycle demonstration and refreshed comments do not bear on the claimed 48/50 SWE-bench run or supply reproducible artifacts, traces, leakage analysis, or an evaluator audit. The case remains an isolated benchmark claim and should stay dormant until auditable evidence appears.
2026-08-24T21:23:11Z
evidence attached: reddit.post.1vxfqhd — The Ox-alpha video appears to be downstream evidence or demonstration relevant to the open benchmark-reproducibility case, though its limited detail weakens the signal.
2026-08-24T20:46:27Z
The refreshed discussion adds no inspectable attribution evidence, independent replication, run artifacts, traces, leakage analysis, or evaluator audit. The 96% result remains an isolated claim, and further provenance discussion does not change the case’s meaning.
2026-08-24T19:55:58Z
The new GLM attribution analysis sharpens the provenance question but lacks first-party confirmation or an inspectable technical artifact, and its own discussion raises conflicting modality evidence. It neither reproduces nor audits the 48/50 run, so the benchmark claim remains isolated and should await auditable artifacts rather than further attribution speculation.
2026-08-24T19:26:37Z
evidence attached: hn.story.49422226 — Independent analysis materially challenges the open case's premise by identifying Ox Alpha as GLM, warranting mention when reassessing the result.
2026-08-24T16:29:14Z
The refreshed discussion remains contradictory model-provenance speculation and adds no comparable run, artifacts, traces, or evaluator audit. The 96% result is still an isolated claim; further attention should wait for auditable replication or a first-party methodology disclosure.
2026-08-24T00:22:55Z
The refreshed discussion and engagement add no matching replication, run artifacts, traces, or evaluator audit; they are further amplification of already-known provenance speculation. The 96% SWE-bench Verified-Mini result remains an isolated claim and should wait for auditable evidence.
2026-08-23T22:25:32Z
The new hands-on report only confirms anecdotal access and ordinary assistant behavior; it contributes no comparable run, traces, artifacts, or evaluator audit. The 96% SWE-bench Verified-Mini result remains an isolated claim, so the case should wait for auditable replication rather than discussion churn.
2026-08-23T22:22:12Z
evidence attached: reddit.post.1vwk6bh — Reports an early hands-on Ox Alpha use, but provides only anecdotal assistant behavior and no benchmark evidence.
2026-08-23T07:22:34Z
The refreshed comments add no inspectable provenance support, comparable replication, run artifacts, traces, or evaluator audit. The 96% result remains an isolated claim, and further discussion churn should not prompt review without auditable evidence.
2026-08-23T02:25:29Z
The comment refresh adds no inspectable provenance evidence, matching replication, run artifacts, traces, or evaluator audit. The 96% result remains an isolated claim, and repetitive discussion should not trigger reassessment absent auditable evidence.
2026-08-23T01:29:01Z
The refreshed provenance discussion adds no inspectable support, first-party identification, matching replication, run artifacts, traces, or evaluator audit. The 96% result remains an isolated claim, and continued comment churn does not merit frequent reassessment.
2026-08-23T00:23:16Z
The new provenance post introduces a possible unreleased-GLM and contamination angle, but provides no inspectable support, first-party confirmation, or audit of the benchmark run. The 96% result remains an isolated claim awaiting artifacts or a comparable independent replication.
2026-08-23T00:22:14Z
evidence attached: reddit.post.1vvrrht — The reported discovery of an unreleased Z.ai GLM model behind Ox Alpha materially contextualizes model provenance and possible benchmark contamination or leakage.
2026-08-22T22:33:01Z
The refreshed discussion still adds no independent matching run, artifacts, traces, or evaluator audit; it remains repetitive provenance and capability speculation. The 96% SWE-bench Verified-Mini result is still an isolated claim and should be revisited only when auditable evidence appears.
2026-08-22T16:30:56Z
The refreshed comments add no independent replication, matching run, artifacts, traces, or evaluator audit; they remain speculative amplification around provenance and adjacent performance. The 96% SWE-bench Verified-Mini claim is still isolated and should only be revisited on auditable evidence.
2026-08-22T13:39:56Z
The latest refresh adds only anecdotal performance criticism, moderation discussion, and minor engagement growth—not a comparable run, artifacts, traces, or evaluator audit. The 96% SWE-bench Verified-Mini result remains an isolated claim and should be revisited only when auditable evidence appears.
2026-08-22T12:29:41Z
The refreshed comments remain contradictory provenance speculation and anecdotal capability impressions, with no matching replication, traces, artifacts, or evaluator audit. The 96% result remains an isolated claim and should now be revisited only when auditable evidence appears.
2026-08-22T11:28:23Z
The refreshed discussion adds no matching replication, artifacts, traces, or evaluator audit, so the 96% result remains an isolated claim. Repeated provenance speculation and anecdotal impressions no longer justify frequent reassessment.
2026-08-22T10:29:53Z
The refreshed comments remain provenance guesses and anecdotal impressions, with no comparable replication, run artifacts, traces, or evaluator audit. The 96% SWE-bench result is still an isolated claim, and discussion churn no longer warrants frequent review.
2026-08-22T09:26:28Z
The comment refresh adds no matching replication, artifacts, traces, or evaluator audit; it is repetitive provenance and capability speculation that does not change the case. The 96% result remains an isolated claim, so further review should wait for auditable evidence rather than discussion churn.
2026-08-22T08:30:11Z
The latest discussion refresh adds no auditable replication, matching run, traces, artifacts, or evaluator review. The 96% result remains an isolated claim, and repeated provenance speculation no longer merits hourly reassessment.
2026-08-22T07:23:34Z
The refreshed comments add only contradictory provenance guesses and anecdotal performance impressions, not a matching replication, traces, artifacts, or evaluator audit. The 96% SWE-bench Verified-Mini result remains an isolated claim with unchanged meaning.
2026-08-22T06:23:16Z
Refreshed comments continue contradictory speculation about Ox Alpha’s identity and adjacent performance without adding a comparable run, traces, artifacts, or evaluator audit. The 96% SWE-bench Verified-Mini result remains an isolated claim awaiting reproducibility evidence.
2026-08-22T05:29:14Z
The added posts broaden speculation about Ox Alpha’s provenance and performance on adjacent benchmarks but provide no independent matching run, traces, artifacts, or evaluator audit. Contradictory identity and capability anecdotes leave the 96% SWE-bench claim isolated and do not change its meaning.
2026-08-22T05:22:27Z
evidence attached: reddit.post.1vv32g8 — The reported benchmark comparison and researcher attribution provide additional, but still unverified, context for the open Ox Alpha evaluation case.
2026-08-22T05:22:27Z
evidence attached: reddit.post.1vv36cs — The screenshot-based report directly bears on Ox Alpha's claimed coding benchmark performance, although it remains weakly evidenced.
2026-08-22T00:24:04Z
The adjacent Multi-Agent Arena claim adds breadth to Ox Alpha’s reported performance but supplies no methodology, artifacts, or independent validation of the SWE-bench run. The 96% result remains an isolated claim awaiting a comparable replication or audit.
2026-08-22T00:22:41Z
evidence attached: hn.story.49395319 — This is additional external performance context for the same Ox Alpha release, although it is a claim rather than independent validation of the open benchmark case.
2026-08-21T22:31:14Z
The refreshed discussion supplies no independent replication, run artifacts, traces, or evaluator audit; it is repetitive amplification and anecdotal skepticism rather than new validation. The 96% result remains an isolated, potentially consequential claim.
2026-08-21T19:30:45Z
The refreshed comments add only speculation and generic benchmark skepticism; no matching replication, traces, artifacts, or evaluator audit changes the evidentiary picture. The 96% result remains a consequential but isolated claim.
2026-08-21T17:54:17Z
A commenter reports a separate roughly 92% result with another model and test setup, which weakly suggests unusually high Mini scores may be broader than Ox Alpha but does not replicate the claimed run. With no matching harness, traces, artifacts, or evaluator audit, the case remains an anomalous single-source result rather than corroborated evidence.
2026-08-21T16:49:58Z
No replication, traces, artifacts, or evaluator audit have appeared; the small engagement increase is repetitive attention rather than corroboration. The 96% result remains an anomalous single-source claim awaiting reproducibility evidence.
2026-08-21T16:29:46Z
grounded: converges/high — The claimed 96% result, coupled with the author’s call for replication and audit, converges directly with Scott’s positions that coding capability must be attri
2026-08-21T16:26:53Z
case created — The reported 48-of-50 result is anomalously strong under a named reproducible harness and warrants prompt independent scrutiny.