Anthropic announced Claude Opus 5 as its new flagship model, positioning it for coding and knowledge work with performance approaching the more expensive Fable 5 at roughly half the token price. Anthropic and secondary reports claim a 30.2% ARC-AGI-3 score, while an ARC Prize results page confirms a dedicated evaluation entry; however, the supplied snippet from that page does not independently expose the aggregate score. Independent benchmarking is beginning to appear, but the provided material is still too thin to establish real-world coding-agent reliability or fully validate Anthropic’s price-performance claims.
2026-07-28T21:24:01Z
The launch-validation window has settled on a durable split verdict: independent evidence supports near-Fable reasoning and coding output, while controlled cost, latency, retry, and reliability results reject the simple half-price value framing. Knowledge-work and ARC-AGI-3 remain unvalidated, but now warrant separate follow-up only if substantive results arrive.
2026-07-28T20:24:37Z
The latest activity adds no inspectable evaluation beyond the established split verdict: independent results support near-Fable reasoning and coding output, while reliability, latency, and cost per successful task undermine the half-price value framing. Engagement is now repetitive amplification, with knowledge-work value and ARC-AGI-3 validity still unresolved.
2026-07-28T19:25:49Z
The behavioral analysis adds qualitative framing for Opus 5’s self-corrective interaction style but no inspectable evaluation that changes the stable split verdict. Near-Fable output remains supported, while workflow reliability and cost per successful task undermine the half-price claim; knowledge-work value and ARC-AGI-3 remain open.
2026-07-28T19:21:14Z
evidence attached: hn.story.49088587 — The behavioral analysis provides qualitative context for whether Opus 5's interaction style affects its practical coding and knowledge-work value.
2026-07-28T18:23:39Z
No fresh evaluative substance changes the stable split verdict: independent results support near-Fable reasoning and coding output, while measured retries, regressions, latency, and cost per successful task undermine the half-price framing. The latest activity is repetitive amplification; knowledge-work value and ARC-AGI-3 validity remain unresolved.
2026-07-28T17:26:01Z
The latest detailed project testimony reinforces the established workflow-reliability gap but adds no independent line beyond MineBench’s split verdict. Near-Fable output capability remains supported, while retries, regressions, latency, and cost per successful task undermine the simple half-price value framing; knowledge-work and ARC-AGI-3 validation remain open.
2026-07-28T16:23:58Z
The detailed project report reinforces that Opus 5’s near-Fable benchmark output can coexist with serious memory, instruction-following, and regression problems in complex coding workflows. It remains anecdotal and mixed in the comments, so it strengthens the established reliability caveat without displacing MineBench’s split verdict or resolving knowledge-work and ARC-AGI-3 claims.
2026-07-28T16:21:40Z
evidence attached: reddit.post.1v92csh — This detailed coding-project report directly contradicts the case hypothesis by describing severe regression, memory, and reliability failures in Opus 5.
2026-07-28T15:25:15Z
No fresh evaluative evidence changes the stable split verdict: independent results support near-Fable reasoning and coding output, while measured retries, latency, and cost per successful task undermine the half-price framing. The remaining activity is repetitive amplification, with knowledge-work value and ARC-AGI-3 validity still unresolved.
2026-07-28T14:27:18Z
The latest changes are engagement churn around already-priced evidence, not a new independent result. The stable split verdict remains: near-Fable reasoning and coding output are corroborated, while measured reliability, latency, and cost per successful task undermine the half-price framing; knowledge-work value and ARC-AGI-3 remain open.
2026-07-28T13:29:54Z
The new activity is repetitive engagement around already-assessed anecdotes and does not change the split verdict: near-Fable reasoning and coding output are supported, while reliability and cost per successful task undermine the half-price framing. Knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-28T12:24:21Z
The latest reports add weak negative evidence that Opus 5’s over-eager, assumption-prone behavior also affects non-coding knowledge work, extending the established workflow caveat beyond coding. They remain small-sample anecdotes, so MineBench’s split verdict still governs and the broader knowledge-work and ARC-AGI-3 claims remain open.
2026-07-28T12:21:42Z
evidence attached: reddit.post.1v8vmed — User comparison of Opus 5 and Fable 5 provides small-sample evidence about differing coding and knowledge-work value, though not a robust evaluation.
2026-07-28T12:21:42Z
evidence attached: reddit.post.1v8w2gf — Anecdotal independent comparison reports Opus 5 underperforming Opus 4.8 on non-coding work, materially challenging its broader value proposition.
2026-07-28T11:23:05Z
The new hands-on comparison adds another favorable but uncontrolled workflow report and does not displace MineBench’s split verdict. Near-Fable output capability remains supported, while reliability, latency, and cost per successful task undermine the half-price framing; knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-28T11:20:59Z
evidence attached: hn.story.49081970 — This is independent hands-on evidence comparing Opus 5 and Fable 5, directly informing the open case about Opus's coding and knowledge-work value.
2026-07-28T10:24:28Z
No new evaluative evidence changes the established split verdict: independent results support near-Fable reasoning and coding output, while measured reliability, latency, and cost per successful task undermine the half-price framing. Current movement is repetitive engagement; knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-28T09:24:41Z
The new low-signal usage anecdote is consistent with Opus 5 being capable and sometimes efficient, but does not alter the controlled split verdict: near-Fable output is supported while reliability, latency, and cost per successful task undermine the half-price framing. Knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-28T09:21:15Z
evidence attached: reddit.post.1v8s1el — This is independent user evidence supporting claims that Opus 5 improves coding speed and accuracy at similar reported usage.
2026-07-28T08:26:27Z
No fresh evaluative evidence changes MineBench’s split verdict: Opus 5 has independently supported near-Fable reasoning and coding output, but reliability, latency, and cost per successful task undermine the half-price framing. The new activity is repetitive engagement, while knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-28T07:25:11Z
The new activity is repetitive engagement around already-assessed demonstrations, not fresh evaluative evidence. MineBench’s split verdict still governs: near-Fable coding output is supported, but reliability, latency, and cost per successful task undermine the half-price framing; knowledge-work value and ARC-AGI-3 remain open.
2026-07-28T06:23:00Z
No new evaluative result changes MineBench’s split verdict: Opus 5 has independently supported near-Fable reasoning and coding output, but reliability, latency, and cost per successful task undermine the half-price framing. The latest movement is repetitive engagement; knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-28T05:21:48Z
No new evaluative substance changes MineBench’s split verdict: Opus 5 has independently supported near-Fable reasoning and coding output, but reliability, latency, and cost per successful task undermine the half-price framing. Current movement is repetitive engagement, while knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-28T04:23:37Z
The latest context-heavy regression report is unsupported testimony and adds only to the established workflow-variability caveat. MineBench’s split verdict still governs: near-Fable coding output is supported, while reliability, latency, and cost per successful task undermine the half-price framing; knowledge-work value and ARC-AGI-3 remain open.
2026-07-28T04:21:06Z
evidence attached: reddit.post.1v8n1j4 — This is a direct user-level report supporting the open hypothesis that Opus 5 may underperform its predecessor on complex context-heavy work.
2026-07-28T03:22:22Z
The new shipped-game optimization and multi-agent 3D builds broaden evidence that Opus 5 is practically capable across agentic coding workflows, but remain uncontrolled demonstrations. They do not alter MineBench’s split verdict: near-Fable output is supported, while reliability, latency, and cost per successful task undermine the half-price framing; knowledge-work value and ARC-AGI-3 remain open.
2026-07-28T03:21:12Z
evidence attached: reddit.post.1v8lj7w — Independent demonstration describes Opus 5 coordinating sub-agents and Blender MCP to produce a playable 3D game, adding substantive evidence about agentic coding breadth.
2026-07-28T03:21:12Z
evidence attached: reddit.post.1v8m095 — Independent use reports Opus 5 building and testing a nontrivial 3D asset, adding qualitative evidence about coding-agent capability.
2026-07-28T03:21:12Z
evidence attached: reddit.post.1v8m8hw — Independent real-world use reports Opus 5 solving a concrete performance problem in a shipped game, modestly supporting its coding value.
2026-07-28T02:23:01Z
No new substantive evaluation changes MineBench’s split verdict: near-Fable coding output is supported, but retries, latency, and cost per successful task contradict the simple half-price framing. The latest movement is repetitive engagement churn, while knowledge-work value and ARC-AGI-3 remain unresolved.
2026-07-28T01:21:30Z
The sustained game-development example is another small, uncontrolled capability anecdote and does not alter MineBench’s split verdict. Near-Fable coding output remains supported, while reliability, latency, and cost per successful task undermine the half-price framing; knowledge-work value and ARC-AGI-3 remain open.
2026-07-28T01:21:05Z
evidence attached: hn.story.49077815 — Independent real-world use of Opus 5 for sustained game development provides a small anecdotal data point on its coding-agent value.
2026-07-28T00:22:00Z
No inspectable SlopCodeBench result or other controlled evidence has arrived to alter MineBench’s split verdict: near-Fable reasoning and coding output are supported, but reliability, latency, and cost per successful task undermine the half-price framing. Knowledge-work value and ARC-AGI-3 validity remain open, while current movement is repetitive engagement churn.
2026-07-27T23:24:13Z
SlopCodeBench signals another independent evaluation path, but the supplied artifact exposes no results, methodology, or cost data, so it cannot change MineBench’s split verdict. Near-Fable coding output remains supported while reliability and cost per successful task undermine the half-price framing; knowledge-work value and ARC-AGI-3 remain open.
2026-07-27T23:21:19Z
evidence attached: hn.story.49076391 — Independent Opus 5 benchmarking directly bears on the open capability-and-value hypothesis.
2026-07-27T22:25:19Z
No new evaluative substance changes the established split verdict: independent results support near-Fable reasoning and coding output, while MineBench’s retries, latency, and cost per successful build contradict the simple half-price value claim. Remaining attention is repetitive amplification; knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-27T21:25:00Z
The new user comparison reinforces the established workflow gap—Fable is often more dependable on long-horizon coding and creative tasks—but remains anecdotal and does not alter MineBench’s controlled split verdict. Near-Fable output capability is supported, while cost per successful task contradicts the half-price framing and knowledge-work and ARC-AGI-3 validation remain open.
2026-07-27T21:21:17Z
evidence attached: reddit.post.1v8cpbr — A comparative user evaluation claims Fable 5 materially outperforms Opus 5 on real coding and creative tasks, challenging the near-parity hypothesis.
2026-07-27T20:24:58Z
No new substantive evidence changes the split verdict: independent results support near-Fable reasoning and coding output, while measured retries, latency, and cost per successful build contradict the simple half-price value claim. Knowledge-work value and ARC-AGI-3 remain unresolved, and current activity is repetitive amplification.
2026-07-27T19:23:17Z
No new evaluative substance changes the split verdict: controlled evidence supports near-Fable reasoning and coding output, while retries, latency, and cost per successful build contradict the half-price value framing. Current activity is repetitive amplification; knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-27T18:24:31Z
No substantive new evaluation changes the split verdict: independent results support near-Fable reasoning and coding output, while measured retries, latency, and cost per successful build refute the simple half-price value framing. The remaining activity is repetitive amplification, with knowledge-work value and ARC-AGI-3 still unresolved.
2026-07-27T17:23:16Z
No new evaluation changes the established split verdict: independent evidence supports near-Fable reasoning and coding output, while measured retries, latency, and cost per successful build undermine the half-price value claim. The latest activity is repetitive amplification, with knowledge-work value and ARC-AGI-3 still unresolved.
2026-07-27T16:28:58Z
The new knowledge-work post is a request for guidance rather than an evaluation, so it adds no evidence beyond MineBench’s split verdict. Near-Fable reasoning and coding output remain supported, while cost-per-success undermines the half-price framing and knowledge-work value and ARC-AGI-3 remain unresolved.
2026-07-27T16:21:54Z
evidence attached: reddit.post.1v848on — A user report that Opus may outperform Fable for long-document and tool-connected knowledge work provides weak independent context for the broader value comparison.
2026-07-27T15:29:47Z
grounded: novel/low — This is broadly Scott’s kind of topic—a frontier model aimed at coding agents and knowledge work—but no wiki or radar hits establish a connection to a specific
2026-07-27T15:29:02Z
The latest activity adds no evaluative substance beyond MineBench’s split verdict: independent evidence supports near-Fable reasoning and coding output, but measured retries, latency, and cost per successful build contradict the simple half-price value framing. Knowledge-work value and ARC-AGI-3 validity remain open, while discussion has settled into repetitive amplification.
2026-07-27T14:28:55Z
The additional comparison is explicitly non-rigorous and merely echoes MineBench’s split verdict: Opus 5 can approach Fable 5’s coding capability, but workflow reliability and iteration costs undermine the simple half-price value claim. It adds no material evidence on knowledge work or ARC-AGI-3.
2026-07-27T14:21:41Z
evidence attached: reddit.post.1v80m4j — This provides an additional, though explicitly non-rigorous, comparison suggesting Opus 5 may lag Fable 5 on complex coding work despite its lower price.
2026-07-27T13:26:44Z
The 200k default context limit adds an operational caveat for long-horizon coding and knowledge work, but appears to concern access configuration rather than a new capability result. It does not change MineBench’s split verdict: near-Fable coding output is supported, while cost, latency, and reliability undermine the half-price value framing.
2026-07-27T13:21:28Z
evidence attached: hn.story.49068610 — The reported 0.2M default context limit materially affects independent assessment of Opus 5's coding and knowledge-work value.
2026-07-27T12:27:16Z
No new controlled evidence changes MineBench’s split verdict: near-Fable coding output is independently supported, but retries, latency, and cost per successful build contradict the half-price value framing. Remaining activity is repetitive amplification, while knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-27T11:29:00Z
No new controlled evidence changes MineBench’s split verdict: near-Fable coding output is supported, but retries, latency, and cost per successful build contradict the simple half-price value claim. The remaining activity is repetitive amplification, while knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-27T10:25:10Z
The latest attachment adds no controlled result beyond MineBench’s split verdict: near-Fable coding output is supported, while retries, latency, and cost per successful build contradict the simple half-price claim. Knowledge-work value and ARC-AGI-3 validity remain open, and subsequent attention is repetitive amplification.
2026-07-27T09:25:14Z
No new controlled evidence changes MineBench’s split verdict: Opus 5 has near-Fable coding output, but retries, latency, and cost per successful build contradict the simple half-price value claim. Current movement is repetitive amplification; knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-27T08:23:15Z
No new controlled result changes MineBench’s split verdict: Opus 5 has independently supported near-Fable coding output, but its retries, latency, and cost per successful build contradict the simple half-price value claim. Remaining activity is repetitive amplification, while knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-27T07:23:50Z
No new evaluative substance changes the split verdict: controlled evidence supports near-Fable coding output but contradicts the half-price framing on latency and cost per successful build. Knowledge-work value and ARC-AGI-3 validity remain open, while current activity is repetitive amplification.
2026-07-27T06:22:59Z
No new evaluative substance changes the split verdict: controlled evidence supports near-Fable coding output but contradicts the half-price framing on latency and cost per successful build. The latest activity is repetitive amplification, while knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-27T05:26:17Z
No new evaluative substance changes the split verdict: controlled evidence supports near-Fable coding output, but contradicts the half-price framing on latency and cost per successful build. Remaining activity is repetitive amplification, while knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-27T04:24:22Z
No new evaluative substance changes the split verdict: controlled evidence supports near-Fable coding output, but contradicts the half-price framing on cost and latency per successful build. Subsequent activity is repetitive amplification, while knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-27T02:23:25Z
No new evaluative substance changes the MineBench repricing: controlled evidence supports near-Fable coding output but contradicts the half-price claim on cost per successful build. Knowledge-work value and ARC-AGI-3 remain unresolved, while subsequent activity is repetitive amplification.
2026-07-27T01:23:10Z
No new evaluative substance extends MineBench; the update is engagement churn around the first controlled coding comparison. Near-Fable coding output is now supported, but the claimed half-price value is contradicted on cost per successful build, while knowledge-work value and ARC-AGI-3 remain unresolved.
2026-07-27T00:22:09Z
No new evaluative substance follows MineBench; the latest movement is engagement churn. Controlled evidence now supports near-Fable coding output but contradicts the simple half-price value claim on cost per successful build, while knowledge-work value and ARC-AGI-3 remain open.
2026-07-26T23:25:02Z
MineBench supplies the first detailed controlled coding-agent comparison: Opus 5 can match or exceed Fable 5 in output capability, but was 78% slower and 64% more expensive across the builds because of retries and invalid schema output. This materially weakens the simple half-price value claim while strengthening near-Fable coding capability; broader knowledge-work value and ARC-AGI-3 validity remain open.
2026-07-26T23:21:13Z
evidence attached: reddit.post.1v7i49g — Independent benchmark (MineBench) shows Opus 5 matching or exceeding Fable 5 on a complex task, with detailed cost and latency data; directly corroborates the case's hypothesis about Opus 5's value.
2026-07-26T22:25:08Z
No new controlled evaluation has arrived since the last look; the newest attachment is another single low-engagement anecdote reinforcing already-known workflow variability (over-eager self-verification, scope creep). Near-Fable general reasoning remains independently supported by SimpleBench, but coding-agent cost-per-task and ARC-AGI-3 validity remain unresolved. Launch-cycle discussion has settled into repetitive amplification; nothing here changes the case's meaning or urgency.
2026-07-26T21:25:05Z
Latest attachment is a single low-engagement anecdote about self-serving harness behavior, consistent with the already-known workflow variability caveat. Near-Fable general reasoning remains independently supported (SimpleBench), but coding-agent cost-per-task and ARC-AGI-3 validity remain unresolved; launch-cycle discussion continues to be repetitive amplification without a controlled evaluation.
2026-07-26T21:21:47Z
evidence attached: reddit.post.1v7fth9 — Anecdotal report of Opus 5 building a harness to approve its own design, corroborating the case's hypothesis about Opus 5's coding-agent capabilities.
2026-07-26T20:24:14Z
The apparent update is repetitive engagement around the already-assessed long-horizon coding complaint, not a new controlled result. Near-Fable general reasoning remains independently supported, while coding-agent cost per successful task and ARC-AGI-3 validity remain unresolved.
2026-07-26T19:25:52Z
The latest movement is engagement churn around the already-assessed long-horizon coding complaint, not a new independent result. Near-Fable general reasoning remains supported, while dependable coding-agent cost per successful task and ARC-AGI-3 validity remain unresolved.
2026-07-26T18:25:05Z
Another long-horizon coding report reinforces that Opus 5’s benchmark-level near-parity may not transfer to dependable agent workflows, particularly around scope control and iteration cost. It remains a low-signal anecdote, so it sharpens the workflow caveat without outweighing independent reasoning results or resolving the broader value claim.
2026-07-26T18:21:18Z
evidence attached: reddit.post.1v7b1u1 — The user reports Opus 5 being materially worse than Fable for long-horizon coding, directly contradicting the near-parity value hypothesis.
2026-07-26T17:28:05Z
The latest post repackages launch benchmarks as validated without supplying new measurements, while its fallback and operational caveats repeat already-known concerns. Near-Fable general reasoning remains independently supported, but coding and knowledge-work cost per successful task and ARC-AGI-3 validity remain unresolved.
2026-07-26T17:21:07Z
evidence attached: reddit.post.1v79tvo — The claimed Opus 5 benchmark and cost advantage directly bear on the open value case, while the alleged fallback and compliance behavior adds operational context.
2026-07-26T16:23:54Z
The latest attachment adds no meaningful independent result beyond the already-assessed mixed usage reports. Near-Fable general reasoning remains supported, but coding and knowledge-work cost per successful task and the ARC-AGI-3 claim remain unresolved; continued launch-cycle churn no longer warrants hourly attention.
2026-07-26T15:25:00Z
The new attachment adds no substantive result beyond already-assessed evidence and appears to be repetitive launch-cycle churn. Near-Fable general reasoning has independent support, but coding and knowledge-work cost per successful task and the ARC-AGI-3 claim remain unresolved pending controlled evaluations.
2026-07-26T14:27:16Z
The latest movement is repetitive engagement rather than new evaluative substance. Independent evidence supports near-Fable general reasoning, but coding and knowledge-work cost per successful task and the ARC-AGI-3 result remain unresolved pending controlled evaluations.
2026-07-26T13:24:13Z
The latest movement is engagement churn rather than a new controlled evaluation. Near-Fable general reasoning remains independently supported, but coding and knowledge-work cost per successful task—and the ARC-AGI-3 result—remain unresolved.
2026-07-26T12:23:57Z
The production report reinforces that Opus 5’s lower token price may not translate into lower cost per successful task because confident omissions require repeated iteration. It is still an unpublished, low-signal comparison, so near-Fable reasoning remains supported while durable coding and knowledge-work value stays unresolved.
2026-07-26T12:20:51Z
evidence attached: reddit.post.1v71up4 — This is an independent production report challenging Opus 5's reliability and cost advantage relative to Fable 5.
2026-07-26T11:24:41Z
The new debugging anecdote is consistent with established practical capability but adds no controlled evidence on coding-agent cost per successful task or ARC-AGI-3. Near-Fable general reasoning remains independently supported, while the broader half-price value claim is still stalled pending evaluations such as MineBench.
2026-07-26T11:21:07Z
evidence attached: reddit.post.1v70fvj — Independent user experience offers weak anecdotal support for Opus 5's practical debugging ability, though it is not a controlled evaluation.
2026-07-26T10:23:51Z
grounded: novel/none — No intersection found in Scott’s wikis, and no radar pages show this release or its claims are already tracked. The supplied material therefore cannot establish
2026-07-26T10:23:11Z
No substantive result extends SimpleBench’s support for near-Fable general reasoning; the new activity is repetitive amplification, while coding-agent cost per successful task and ARC-AGI-3 validity remain unresolved. The cached grounding still treats the release as unconfirmed and is now materially stale.
2026-07-26T09:23:55Z
No new substantive result extends the SimpleBench evidence; current movement is repetitive engagement rather than validation. Near-Fable general reasoning has independent support, but coding-agent cost per successful task and ARC-AGI-3 validity remain unresolved pending controlled evaluations.
2026-07-26T08:21:56Z
No new substantive evaluation follows the already-priced SimpleBench result; the latest movement is engagement churn. Near-Fable general reasoning now has independent support, but coding-agent price-performance and ARC-AGI-3 validity remain unresolved.
2026-07-26T07:22:51Z
SimpleBench independently places Opus 5 within 1.3 points of Fable 5, strengthening the case from generic frontier capability to measured near-Fable reasoning performance. It still does not validate coding-agent performance, cost per successful task, or the reported ARC-AGI-3 result, so the full value claim remains open.
2026-07-26T07:20:55Z
evidence attached: reddit.post.1v6w0c5 — An independent benchmark places Opus 5 just behind Fable 5 and well ahead of prior Opus versions, providing useful capability context for the value comparison.
2026-07-26T06:23:10Z
The latest activity adds no completed controlled evaluation and remains repetitive amplification of mixed, uncontrolled usage reports. Opus 5’s frontier capability is corroborated, but durable near-Fable price-performance and the ARC-AGI-3 result remain unsettled pending results such as MineBench.
2026-07-26T05:21:58Z
No completed controlled evaluation has arrived; the latest movement is repetitive engagement around previously assessed anecdotes. Frontier capability remains corroborated, but durable near-Fable price-performance and ARC-AGI-3 validity remain unsettled pending results such as MineBench.
2026-07-26T04:21:50Z
The apparent movement is engagement churn around the already-assessed Blender anecdote, not new independent validation. Frontier capability remains corroborated, while durable near-Fable price-performance and ARC-AGI-3 validity remain unsettled pending controlled results such as MineBench.
2026-07-26T03:22:09Z
The Blender workflow adds another favorable but uncontrolled implementation anecdote, reinforcing practical capability without validating durable near-Fable price-performance or the ARC-AGI-3 result. The case remains stalled pending controlled independent evaluations such as MineBench.
2026-07-26T03:21:09Z
evidence attached: reddit.post.1v6pvby — Independent use reports strong Opus 5 performance on a nontrivial Blender workflow, adding capability evidence to the model-value case.
2026-07-26T02:22:08Z
The latest movement is engagement churn around already-assessed anecdotes, with no completed controlled evaluation of price-performance or ARC-AGI-3. Frontier capability remains corroborated, but the stronger near-Fable-at-half-price claim is stalled pending results such as MineBench.
2026-07-26T01:23:46Z
The latest additions amount to engagement churn around already-assessed anecdotes, not the controlled evaluation needed to validate near-Fable price-performance or ARC-AGI-3. Frontier capability remains corroborated, but the stronger value claim is stalled pending results such as MineBench.
2026-07-26T00:23:57Z
No substantive evaluation has arrived beyond the already-assessed lightweight stress test; the latest movement is repetitive engagement rather than new validation. Frontier capability remains corroborated, while durable near-Fable price-performance and the ARC-AGI-3 result remain unsettled pending controlled results.
2026-07-25T23:25:37Z
The lightweight custom stress test adds another uncontrolled comparison but no credible validation of durable near-Fable price-performance or the ARC-AGI-3 result. The case remains stalled at established frontier capability, with discussion now largely repetitive amplification.
2026-07-25T23:21:05Z
evidence attached: reddit.post.1v6lw34 — This is an independent, albeit lightweight, evaluation bearing on whether Claude Opus 5 has a durable frontier-reasoning advantage.
2026-07-25T22:29:27Z
Another small game demo broadens anecdotal support for Opus 5’s coding ability but remains uncontrolled and does not validate durable near-Fable price-performance or ARC-AGI-3. The discussion is repetitive amplification, so the case remains stalled pending controlled independent results.
2026-07-25T22:21:27Z
evidence attached: reddit.post.1v6l3nx — A small independent demonstration supports Opus 5's ability to produce substantial working software, though it is anecdotal.
2026-07-25T21:24:40Z
The system card adds authoritative context on Opus 5’s intended capabilities and operating caveats, but remains first-party evidence and does not validate the near-Fable price-performance or ARC-AGI-3 claims. The case is still waiting on completed controlled evaluations rather than further launch commentary.
2026-07-25T21:20:56Z
evidence attached: hn.story.49051539 — The Opus 5 system card materially contextualizes the model's claimed capability and safety profile, though it is not independent validation.
2026-07-25T20:26:11Z
No completed independent evaluation has arrived; the apparent movement remains engagement around mixed anecdotes and the promised MineBench test. Frontier-level capability is corroborated, but near-Fable price-performance and ARC-AGI-3 validity remain unsettled.
2026-07-25T19:24:03Z
The latest movement is engagement and comment churn around already-known anecdotes, not a completed independent evaluation. Frontier-level capability remains corroborated, while near-Fable price-performance and ARC-AGI-3 validity remain unsettled pending controlled results such as MineBench.
2026-07-25T18:29:46Z
The new knowledge-work test adds a concrete but very weak negative datapoint, while MineBench is only a promised evaluation rather than a result. The case remains stalled: frontier capability is corroborated, but near-Fable price-performance and ARC-AGI-3 validity are still unsettled.
2026-07-25T18:21:46Z
evidence attached: reddit.post.1v6eyb4 — The announced MineBench evaluation will provide an independent data point on whether Opus 5 has a durable coding-agent advantage.
2026-07-25T18:21:46Z
evidence attached: reddit.post.1v6ehpd — Independent Opus 5 testing provides relevant evidence on knowledge-work quality, including notable factual and selection errors.
2026-07-25T17:24:42Z
The new promotional one-shot demo is another uncontrolled coding anecdote, not independent validation of durable near-Fable price-performance or the ARC-AGI-3 result. Its modest engagement adds attention but no substantive change, so the case remains stalled at frontier-level capability corroboration.
2026-07-25T17:21:22Z
evidence attached: reddit.post.1v6dvey — Anecdotal independent coding demo modestly supports Opus 5's coding-agent value case, though the promotional presentation makes it weak evidence.
2026-07-25T16:25:03Z
The latest material remains conflicting, uncontrolled usage testimony rather than a controlled independent evaluation. It reinforces that Opus 5 is frontier-capable but that its price-performance varies by workload, harness, effort setting, and account behavior, leaving the near-Fable and ARC-AGI-3 claims unresolved.
2026-07-25T15:24:16Z
Conflicting token-efficiency and production-use reports increasingly suggest that Opus 5’s value depends on workload, harness, effort setting, and account behavior rather than delivering a dependable half-price near-Fable advantage. These remain uncontrolled anecdotes, so they sharpen the variability caveat without validating or disproving the core claim.
2026-07-25T14:27:39Z
Fresh hands-on reports now conflict even on token efficiency, reinforcing that Opus 5’s value is highly sensitive to account, effort setting, harness, and workload rather than a durable half-price advantage. Without published measurements or controlled comparisons, this broadens the variability caveat but does not advance the core validation case.
2026-07-25T14:21:18Z
evidence attached: reddit.post.1v68qnj — A substantive user claims inconsistent Opus 5 and Fable quality across thousands of application requests, materially challenging the durability of the value hypothesis despite lacking published measurements.
2026-07-25T14:21:18Z
evidence attached: reddit.post.1v6973n — Independent user experience reports lower token usage and strong results from Opus 5, providing weak but relevant positive evidence for its coding value relative to Fable.
2026-07-25T13:25:03Z
The claimed difficult real-world fix is potentially valuable but remains an unverified, zero-engagement anecdote with a promised follow-up rather than inspectable evidence. It does not change the established frontier-level assessment or resolve near-Fable price-performance and ARC-AGI-3 validity.
2026-07-25T13:21:02Z
evidence attached: reddit.post.1v681u7 — This is a concrete but highly anecdotal independent coding success for Opus 5 relative to competing models.
2026-07-25T12:23:27Z
The enterprise-work report adds another mixed anecdote, while the benchmaxxing discussion raises a plausible but weakly evidenced challenge to the ARC-AGI-3 headline. Neither changes the established frontier-level assessment or supplies the controlled price-performance and benchmark validation the stronger claim requires.
2026-07-25T12:20:57Z
evidence attached: reddit.post.1v66o8k — The reported ARC-AGI score bears directly on the open case's claim about Opus 5's capability, though the screenshot-only evidence is weak.
2026-07-25T12:20:57Z
evidence attached: reddit.post.1v6657a — A user reports Opus 5 matching the premium model for enterprise architecture work, providing weak independent support for its value proposition.
2026-07-25T11:24:34Z
The small RPG prototype adds another favorable but uncontrolled coding anecdote and does not change the case’s meaning. Frontier-level capability is established, while near-Fable price-performance and ARC-AGI-3 validity still await controlled independent evaluation.
2026-07-25T11:20:53Z
evidence attached: hn.story.49046504 — This is weak but relevant independent use suggesting Opus 5 may deliver strong coding-agent value, though the evidence is only a small anecdotal prototype.
2026-07-25T10:24:24Z
The latest additions do not advance the case beyond established frontier-level capability; activity remains repetitive launch amplification and deployment context. Controlled real-codebase, cost-per-task, and ARC-AGI-3 validation are still missing, leaving the near-Fable-at-half-price claim unsettled.
2026-07-25T09:29:49Z
Latest addition (Microsoft Foundry availability) is minor deployment-context news, not evaluative evidence. The case remains stalled at frontier-level capability corroboration with mixed anecdotal signals; near-Fable price-performance and ARC-AGI-3 validity still lack controlled independent evaluation, and launch-cycle discussion has become largely repetitive.
2026-07-25T09:21:07Z
evidence attached: reddit.post.1v63d9s — Corroborates Claude Opus 5 availability on another platform (Microsoft Foundry), supporting the case about its value and deployment.
2026-07-25T08:24:29Z
The latest activity adds no independent evaluation beyond the already-known effort-tuning and model-identification caveats; it is mostly repetitive launch amplification. Frontier-level capability remains corroborated, but near-Fable price-performance and the ARC-AGI-3 claim remain unresolved.
2026-07-25T07:23:13Z
The new evidence sharpens the practical caveat: Opus 5’s coding value depends on effort tuning and clean model identification, since higher effort can trigger over-refactoring and fallback behavior can muddy comparisons. This further weakens benchmark-to-workflow generalization without resolving the near-Fable price-performance or ARC-AGI-3 claims.
2026-07-25T07:20:55Z
evidence attached: reddit.post.1v60pga — Independent coding observations that higher effort reduces scores and increases over-refactoring materially qualify Opus 5's reliability and value claims.
2026-07-25T06:24:02Z
No substantive new evaluation has arrived; the apparent change is repetitive engagement around existing launch claims and uncontrolled demos. Frontier-level capability remains corroborated, while near-Fable price-performance and ARC-AGI-3 validity remain unresolved.
2026-07-25T05:22:58Z
No substantive independent evaluation has arrived; the new activity is repetitive launch amplification rather than evidence resolving price-performance or ARC-AGI-3 validity. Frontier-level capability remains corroborated, but the stronger near-Fable-at-half-price claim is still unsettled.
2026-07-25T04:24:42Z
The context-window concern appears to be rollout or model-selection confusion, while the usage-planning post adds no measured comparison. These additions do not advance validation of near-Fable price-performance or ARC-AGI-3, and the launch discussion is now mostly repetitive amplification.
2026-07-25T04:20:53Z
evidence attached: reddit.post.1v5vt98 — This provides independent albeit anecdotal evidence about how users may divide work between Opus 5 and Fable 5 at the reported price difference.
2026-07-25T04:20:53Z
evidence attached: reddit.post.1v5x398 — The reported context-window discrepancy materially contextualizes Opus 5's practical value, though the evidence is only a low-engagement user observation.
2026-07-25T03:22:25Z
The added activity remains launch amplification and uncontrolled coding demos, broadening anecdotal support without materially changing the case. Frontier-level capability is corroborated, but near-Fable price-performance and the ARC-AGI-3 result still need controlled independent validation.
2026-07-25T02:21:29Z
Several autonomous coding demos broaden the practical evidence that Opus 5 can produce strong long-horizon results, and one toy comparison supports lower cost than Fable 5. The tasks remain uncontrolled and subjective, so they do not validate the general near-Fable value claim or the ARC-AGI-3 result.
2026-07-25T02:21:04Z
evidence attached: reddit.post.1v5ujes — A comparative 3D physics coding result adds weak external evidence about Opus 5’s capability and cost relative to Fable 5 and competing models.
2026-07-25T02:21:04Z
evidence attached: reddit.post.1v5v1in — A user reports Opus 5 producing a functional multi-room interactive 3D environment in one prompt, providing anecdotal coding-agent evidence.
2026-07-25T02:21:04Z
evidence attached: reddit.post.1v5v71s — A user reports a substantial autonomous web-building result from Opus 5, useful but weak anecdotal evidence of coding-agent capability.
2026-07-25T01:21:21Z
The latest additions and engagement are repetitive launch amplification rather than new validation. Independent evidence still supports frontier-level capability, but the near-Fable price-performance and ARC-AGI-3 claims remain unsettled pending controlled coding and cost-per-task evaluations.
2026-07-25T00:21:34Z
The recap confirms rollout and access context but adds no substantive capability or cost-per-task validation. The case remains frontier-level but mixed, with the near-Fable value and ARC-AGI-3 claims still awaiting controlled independent evaluation.
2026-07-25T00:21:09Z
evidence attached: reddit.post.1v5rxo2 — The recap adds product availability, pricing, and usage-limit context relevant to evaluating Opus 5's practical value and adoption.
2026-07-24T23:21:55Z
Artificial Analysis now independently supports Opus 5 as a frontier-level model, while separate hands-on reports provide a second, though mixed, line of practical evidence. The stronger near-Fable-at-half-price and ARC-AGI-3 claims remain unsettled because coding usability is inconsistent and cost-per-task evidence complicates the launch framing.
2026-07-24T23:20:58Z
evidence attached: hn.story.49040741 — Independent leaderboard placement is corroborating evidence that Opus 5 has frontier-level capability, relevant to whether its premium value claims hold.
2026-07-24T23:20:58Z
evidence attached: reddit.post.1v5qvgj — A hands-on comparison of Opus 5 on a demanding coding task provides weak but relevant independent evidence about its practical capability.
2026-07-24T23:20:58Z
evidence attached: reddit.post.1v5r0gz — A user reports current Opus 5 coding-agent usability problems and reverting to an older model, modestly contextualizing value and reliability claims.
2026-07-24T22:26:27Z
The first direct Opus 5–Fable 5 comparison is concrete but uncontrolled and yields a mixed result: Opus appears visually impressive while possibly sacrificing task fidelity. It does not materially validate the broad capability, ARC-AGI-3, or half-price value claims, and launch amplification is becoming repetitive.
2026-07-24T22:21:17Z
evidence attached: reddit.post.1v5p34g — Direct user comparison of Opus 5 and Fable 5 offers weak anecdotal evidence relevant to Opus's claimed coding value.
2026-07-24T21:22:25Z
The latest material suggests Opus 5’s gains may depend heavily on self-verification and verifiable-task settings, while launch-day cost reporting further complicates the half-price framing. These are useful qualifications but still do not provide the independent real-codebase or ARC-AGI-3 validation needed for corroboration.
2026-07-24T21:21:13Z
evidence attached: hn.story.49041500 — Independent launch-day pricing and cost-per-task reporting directly bears on whether Opus 5's coding gains justify its premium cost.
2026-07-24T21:21:13Z
evidence attached: hn.story.49041501 — The leaked Opus 5 system prompt provides contextual evidence about the model and its intended behavior, though it is weakly sourced.
2026-07-24T21:21:13Z
evidence attached: reddit.post.1v5mxfl — A hands-on analysis materially contextualizes whether Opus 5's reported gains come from self-verification workflows rather than broadly transferable capability.
2026-07-24T21:21:13Z
evidence attached: reddit.post.1v5o56l — This provides user-level context on Opus 5's domain capability, safeguards, and limits relevant to evaluating its practical premium value.
2026-07-24T20:24:13Z
Early hands-on reports now suggest a plausible long-horizon and low-effort cost advantage, but the ARC replay shows that the headline aggregate may conceal weak, expensive task behavior. The evidence remains anecdotal or derivative rather than the independent real-codebase validation needed to establish the near-Fable value claim.
2026-07-24T20:21:35Z
evidence attached: hn.story.49040691 — The reported Opus-versus-GPT speed comparison materially contextualizes whether Anthropic models offer a performance or cost advantage, though the version mismatch weakens it.
2026-07-24T20:21:35Z
evidence attached: reddit.post.1v5ll4z — This user-reported benchmark summary supports Opus 5’s value claim while the cited ARC replay supplies potentially important counterevidence.
2026-07-24T20:21:35Z
evidence attached: reddit.post.1v5le69 — A relatively detailed user report supports Opus 5’s claimed long-horizon and cost-adjusted coding advantage, though it remains anecdotal.
2026-07-24T19:24:54Z
Early usage reports are now mixed, and cost-per-task discussion weakens the simple “half-price near-Fable” framing. The additions remain anecdotal or derivative, so the core capability, ARC-AGI-3, and value claims still await substantive independent evaluation.
2026-07-24T19:21:15Z
evidence attached: reddit.post.1v5jkml — Current pricing evidence directly bears on whether Opus 5 delivers enough capability per dollar to justify its value proposition.
2026-07-24T19:21:15Z
evidence attached: reddit.post.1v5jokp — The reported perfect IMO score is a relevant capability datapoint for evaluating Opus 5, though independently unverified.
2026-07-24T19:21:15Z
evidence attached: reddit.post.1v5jxqd — This user report is weak but provides anecdotal negative evidence against Opus 5's claimed durable coding advantage.
2026-07-24T19:21:15Z
evidence attached: reddit.post.1v5jlpo — The reported ARC-AGI-3 benchmark result bears on Opus 5's claimed frontier capability, though the post provides little evidence.
2026-07-24T19:21:15Z
evidence attached: reddit.post.1v5jrx0 — A user report of confident failures on a difficult task provides negative evidence relevant to Opus 5's claimed quality.
2026-07-24T19:21:15Z
evidence attached: reddit.post.1v5k40x — This is a launch confirmation for the already-open Opus 5 capability-validation episode, though largely redundant coverage.
2026-07-24T19:21:15Z
evidence attached: reddit.post.1v5k4w1 — The reported accuracy-forward behavior is relevant contextual evidence for evaluating Opus 5's practical coding-agent tradeoffs.
2026-07-24T19:21:15Z
evidence attached: reddit.post.1v5l56y — Early user report on Opus 5 provides weak initial evidence about its coding quality and model-to-model teaching behavior.
2026-07-24T18:25:12Z
grounded: novel/none — No intersection found: there are no Scott wiki hits connecting the unconfirmed Opus 5 claims to his positions or projects, and no radar hits showing this develo
2026-07-24T18:24:32Z
The launch itself is now substantiated and an Artificial Analysis listing opens the first independent validation path, but the near-Fable value and ARC-AGI-3 claims remain dominated by vendor material and repetitive amplification rather than detailed independent or real-codebase results.
2026-07-24T18:21:45Z
evidence attached: hn.story.49038856 — The official Opus 5 model documentation provides first-party capability and pricing details relevant to evaluating its coding and knowledge-work value.
2026-07-24T18:21:45Z
evidence attached: hn.story.49039047 — Independent Artificial Analysis results directly bear on whether Opus 5 delivers near-Fable capability at substantially lower cost.
2026-07-24T18:21:45Z
evidence attached: reddit.post.1v5hrbc — Release details and claimed benchmark advantage materially contextualize the open question of Opus 5's coding capability and price-performance.
2026-07-24T18:21:45Z
evidence attached: reddit.post.1v5hujd — User discussion bears on whether Opus 5 is genuinely the leading coding model, though it provides no independent evaluation.
2026-07-24T18:21:45Z
evidence attached: reddit.post.1v5idua — Directly motivates real-codebase comparison of Opus 5 against Fable 5, a key test of the open value hypothesis.
2026-07-24T18:21:45Z
evidence attached: reddit.post.1v5ixn6 — Anecdotal comparison favoring Opus over Fable is a weak but directly relevant datapoint for the open value-validation case.
2026-07-24T17:24:10Z
grounded: novel/none — No intersection found: there are no Scott wiki or radar hits connecting this purported release to his existing positions, projects, or tracked cases. The suppli
2026-07-24T17:23:31Z
origin walked (codex/luna, conf 0.99): anchor reddit.post.1v5h6o9 -> echo.blog.e75df1016a by Anthropic
2026-07-24T17:22:55Z
case created — A first-party frontier-model release is already producing multiple benchmark and rollout observations that warrant rapid independent validation.