2026-10-11 17:10 UTC

Reddit builder KangarooAnxious9394's pre-registered experiments claim escalating a cheap coding agent (Haiku) to a strong model (Sonnet) only when it repeats mistakes added 7 successes in 63 runs at ~1.3x cost, beating an always-on advisor; independent replication on unseen repos would establish repeat-mistake-triggered escalation plus fact-reporting verification as a standard cheap-agent harness pattern.

state: seedheat: lowuncertainty: mediumconvergesscott: highagent-harnesses model-escalation cost-optimization

What is this?

A Reddit user (KangarooAnxious9394) cross-posted to r/codereview and r/LocalLLaMA a writeup titled 'I tested 20+ ways to make a cheap coding model act like an expensive one,' reporting that the winning intervention was selective escalation: a stronger model (Haiku β†’ Sonnet) that only intervenes when the cheap agent repeats a mistake, adding +7 successes across 63 runs at roughly 1.3x cost, reportedly beating an always-on advisor setup. The snippets confirm the headline numbers only via the poster's own text β€” no independent replication, no details on methodology, task suite, or repos β€” so the claims are currently single-source self-reports.

Why it matters to Scott

An independent builder arrives at Scott's Model Barbell territory with a refinement his canon doesn't yet hold: mistake-recurrence as the escalation trigger (vs. budget, confidence, or task-type triggers he documents in cost-tiered routing and risk-based triage), plus a pre-registered beats-always-on-advisor number (+7/63 at ~1.3x) that is dated ammunition for his cheap-front-door economics. It is directly actionable β€” the trigger would slot into his LiteLLM tier chain and `ask` harness, and he already owns the fixtures (trace-backed comparison, brief A/B) to replicate it β€” though the note's caveat stands: single-source self-report, no independent replication yet, and the modest delta is itself evidence for his Scout–Senior claim that escalation routing alone is adjacent, not sufficient.
ip:concept.model-barbellip:concept.risk-based-triagedev:concept.cost-tiered-llm-routingdev:concept.cheap-model-front-doordev:concept.trace-backed-agent-comparisondev:project.askdev:technology.litellmradar:concept.model-routingradar:concept.inference-economicsradar:concept.agent-memoryradar:learned-router-task-identityradar:dynamic-model-switching-evaluationradar:haiku-sonnet-task-length-gapradar:multi-model-orchestrator-worker-agents
queries asked of Scott's wikis
  • escalation routing cheap-to-strong model handoff in agent harnesses
  • repeat-mistake detection as agent memory signal for triggering stronger models
  • Haiku vs Sonnet cost-per-solved-task coding agent benchmarks
  • selective critic/advisor pattern vs always-on verifier cost economics
  • pre-registered experiment methodology for agent harness comparisons

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 139h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

10-05 20:39⭐ origin directly observedI tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't
KangarooAnxious9394 on r/LocalLLaMA
β€”
10-05 20:39amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wyjmzv
KangarooAnxious9394
peak 0 Β· 9 comments Β· 100% of case engagement
10-05 23:20our radar first saw it Β· +2.7hdiscovery anchor: reddit.post.1wyjmzvβ€”
pace: p47 vs 1247 stories at the 96h mark (now 139h old) β€” ahead of aws-agentcore-credential-exposure-containment-failure (1.1x), behind asksary-liveloop-stateful-editing (0.9x)

Evidence (1) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐I tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't
LocalLLaMA
KangarooAnxious939409

Interpretation history

Decision trace