EvalRaccoonDev reports that Haiku 4.5 ties Sonnet 4.6 on short tasks but trails it 42.0% to 85.6% overall in a linked same-harness evaluation, suggesting short coding benchmarks understate the reliability gap when selecting models for longer agent workflows.
state: seedheat: mediumuncertainty: mediumcontradictsscott: highagent-evaluation long-horizon-agents coding-agentsEvalRaccoonDev
What is this?
EvalRaccoonDev (apparently an independent evaluator on X/Twitter) reports a same-harness comparison in which Claude Haiku 4.5 ties Claude Sonnet 4.6 on short coding tasks but trails it 42.0% to 85.6% on their internal long-horizon agent tasks, concluding that short benchmarks like SWE-bench understate the reliability gap. The web snippets corroborate the general pattern from public benchmarks โ SWE-bench Verified shows a modest ~6-point gap (73.3% vs 79.6%), while third-party write-ups report the gap widening on multi-step reasoning (10+ points) โ but none of the supplied snippets confirm the specific 42.0%/85.6% internal numbers, the harness, or who EvalRaccoonDev is, so the headline claim rests on unverified original reporting. Anthropic's own launch framing of Haiku 4.5 as matching Sonnet 4 on coding is noted in the snippets as a source of the confusion the eval pushes against.
Why it matters to Scott
A same-harness eval showing Haiku 4.5 tying Sonnet on short tasks but collapsing 42%โ85.6% on long-horizon agent work directly challenges the cheap-model-default assumption wired into Scott's cost-tiered routing (ask's --fast tier, Synthetic Futures' cheap-model repair gates): short-benchmark parity would route long-horizon agent workflows to a model that fails them. It simultaneously validates his trace-backed same-harness comparison methodology and echoes the radar's dynamic-model-switching finding that static/short evaluations mis-rank models for real agent workloads.
dev:concept.cost-tiered-llm-routingdev:concept.task-aware-model-routingdev:concept.trace-backed-agent-comparisondev:project.askip:framework.12-factor-agents-frameworkradar:concept.model-routingradar:concept.agent-evaluationradar:concept.agent-benchmarksradar:concept.long-horizon-agentsradar:dynamic-model-switching-evaluationradar:ai-benchmark-saturation-distortion
queries asked of Scott's wikis
- long-horizon agent evaluation methodology internal harness
- model routing cost-quality tradeoff coding agent Haiku Sonnet
- error compounding multi-step agent reliability degradation
- SWE-bench short benchmark validity selecting production models
- cheap model escalation fallback strategy agent workflows
- agent eval harness project own task suite
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 459h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p54 vs 1032 stories at the 336h mark (now 459h old) โ ahead of agentdrive-persistent-shared-storage (1.1x), behind aws-project-spend-limits (0.9x)
Evidence (1) โ โญ canonical anchor
Interpretation history
2026-09-23T19:37:01Z
grounded: contradicts/high โ A same-harness eval showing Haiku 4.5 tying Sonnet on short tasks but collapsing 42%โ85.6% on long-horizon agent work directly challenges the cheap-model-defaul
2026-09-22T14:24:52Z
case created โ The concrete comparison links tasks and harness code, although the truncated account leaves sample sizes and methodological controls unresolved.
Decision trace
- 09-26 00:30review_screenjev screen: no material development (noul=0.05)
- 09-24 05:37groundA same-harness eval showing Haiku 4.5 tying Sonnet on short tasks but collapsing 42%โ85.6% on long-horizon agent work directly challenges the cheap-model-default assumption wired into Scott's cos
- 09-24 04:15createThe concrete comparison links tasks and harness code, although the truncated account leaves sample sizes and methodological controls unresolved.