2026-10-11 17:13 UTC

EvalRaccoonDev reports that Haiku 4.5 ties Sonnet 4.6 on short tasks but trails it 42.0% to 85.6% overall in a linked same-harness evaluation, suggesting short coding benchmarks understate the reliability gap when selecting models for longer agent workflows.

state: seedheat: mediumuncertainty: mediumcontradictsscott: highagent-evaluation long-horizon-agents coding-agentsEvalRaccoonDev

What is this?

EvalRaccoonDev (apparently an independent evaluator on X/Twitter) reports a same-harness comparison in which Claude Haiku 4.5 ties Claude Sonnet 4.6 on short coding tasks but trails it 42.0% to 85.6% on their internal long-horizon agent tasks, concluding that short benchmarks like SWE-bench understate the reliability gap. The web snippets corroborate the general pattern from public benchmarks โ€” SWE-bench Verified shows a modest ~6-point gap (73.3% vs 79.6%), while third-party write-ups report the gap widening on multi-step reasoning (10+ points) โ€” but none of the supplied snippets confirm the specific 42.0%/85.6% internal numbers, the harness, or who EvalRaccoonDev is, so the headline claim rests on unverified original reporting. Anthropic's own launch framing of Haiku 4.5 as matching Sonnet 4 on coding is noted in the snippets as a source of the confusion the eval pushes against.

Why it matters to Scott

A same-harness eval showing Haiku 4.5 tying Sonnet on short tasks but collapsing 42%โ†’85.6% on long-horizon agent work directly challenges the cheap-model-default assumption wired into Scott's cost-tiered routing (ask's --fast tier, Synthetic Futures' cheap-model repair gates): short-benchmark parity would route long-horizon agent workflows to a model that fails them. It simultaneously validates his trace-backed same-harness comparison methodology and echoes the radar's dynamic-model-switching finding that static/short evaluations mis-rank models for real agent workloads.
dev:concept.cost-tiered-llm-routingdev:concept.task-aware-model-routingdev:concept.trace-backed-agent-comparisondev:project.askip:framework.12-factor-agents-frameworkradar:concept.model-routingradar:concept.agent-evaluationradar:concept.agent-benchmarksradar:concept.long-horizon-agentsradar:dynamic-model-switching-evaluationradar:ai-benchmark-saturation-distortion
queries asked of Scott's wikis
  • long-horizon agent evaluation methodology internal harness
  • model routing cost-quality tradeoff coding agent Haiku Sonnet
  • error compounding multi-step agent reliability degradation
  • SWE-bench short benchmark validity selecting production models
  • cheap model escalation fallback strategy agent workflows
  • agent eval harness project own task suite

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 459h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-22 12:49โญ origin directly observedHaiku is 6 points behind Sonnet on SWE-bench and 44 points behind on our internal agent tasks. But on short tasks they tie.
EvalRaccoonDev on r/ClaudeAI
โ€”
09-22 12:49amplified on r/ClaudeAI ๐Ÿ‘‘reddit.post.1wn8m15
EvalRaccoonDev
peak 9 ยท 10 comments ยท 101% of case engagement
09-22 13:20our radar first saw it ยท +0.5hdiscovery anchor: reddit.post.1wn8m15โ€”
pace: p54 vs 1032 stories at the 336h mark (now 459h old) โ€” ahead of agentdrive-persistent-shared-storage (1.1x), behind aws-project-spend-limits (0.9x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญHaiku is 6 points behind Sonnet on SWE-bench and 44 points behind on our internal agent tasks. But on short tasks they tie.
ClaudeAI
EvalRaccoonDev910

Interpretation history

Decision trace