bfeeny's controlled experiment claims a learned cheap-vs-expert router scoring 0.84 held-out AUC still scores 0.838 when within-task labels are shuffled — it learned task identity, not difficulty — leaving output-based deferral, not learned routing, as the practical cost-saving mechanism until replication finds genuine difficulty signal.
state: corroboratedheat: lowuncertainty: mediumconvergesscott: highllm-routing inference-economics evaluation-methodology
What is this?
This case tracks a negative-result replication claim by Brian Feeny (bfeeny): his 21 Sep 2026 post 'Replicating RouteLLM on Amazon Bedrock' reports that a learned cheap-vs-expert router with 0.84 held-out AUC still scores ~0.838 when cheap/expert labels are shuffled within each task — i.e. the router learned which task (hence which model) produced a query, not per-query difficulty — so the practical cost-saving mechanism is output-based deferral (run the cheap model, escalate on failure) rather than learned routing. A second, methodologically unrelated line (KangarooAnxious9394, Oct 2026, pre-registered Haiku-cheap/Sonnet-strong experiments on real repo commits) independently corroborates the behavioral half: struggle-triggered escalation netted +7 successes at ~1.3x cost, with fact-reporting beating advice. Strong caveats: the supplied web results contain nothing on this work (they return unrelated MoE-router and generic-ML papers), the within-task shuffle control is single-author and unreplicated by third parties, both posts drew near-zero community scrutiny (score 0, upvote ratios 0.17–0.2), and the canonical blog anchor is an echo-reconstructed testimony source — the numbers rest entirely on the authors' own reporting.
Why it matters to Scott
Independent outsiders arrive, via a confound-eliminated negative result plus a methodologically unrelated pre-registered corroboration, exactly where Scott's canon already stands: the Scout-and-the-Senior demotion of generic small/large-model routing and his default-cheap-with-fallback designs (cost-tiered routing, metacognitive failure escalation, the propose–finalise gate) all bet on behavioral deferral over learned difficulty — this is the dated receipt for that bet. The within-task shuffle control is itself the reusable contribution: a cheap falsification test to demand before crediting any learned/adaptive router's claimed savings, though it remains single-author and unreplicated, so it arms his skepticism rather than settles the question.
dev:concept.cost-tiered-llm-routingdev:concept.task-aware-model-routingdev:concept.metacognitive-resolution-controldev:concept.propose-finalise-gateradar:ramp-thompson-sampling-model-routingradar:dynamic-model-switching-evaluationradar:world-model-optimizer-agent-routingradar:nemo-switchyard-llm-routerradar:yhahn-agent-escalation-ergonomics
queries asked of Scott's wikis
- scout-senior split small-model large-model routing demotion
- model barbell default-cheap strong-model fallback design
- agent escalation on failure harness struggle detection
- evaluation confounds shortcut learning benchmark leakage
- inference economics cost-saving claims skepticism
- router difficulty signal vs output-based deferral
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p50momentum: steady2 platformsage 506h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p9 vs 1032 stories at the 336h mark (now 506h old) — behind addom-local-coding-harness (0.5x)
Evidence (3) — ⭐ canonical anchor
Interpretation history
2026-10-06T02:10:33Z
grounded: converges/high — Independent outsiders arrive, via a confound-eliminated negative result plus a methodologically unrelated pre-registered corroboration, exactly where Scott's ca
2026-10-06T02:02:49Z
The thesis moved from single-author negative result to independently corroborated: KangarooAnxious9394's pre-registered Haiku/Sonnet repo-commit experiments found struggle-triggered escalation delivers real savings (+7 successes at ~1.3x cost), a methodologically unrelated second line converging on 'behavioral deferral beats learned difficulty routing' — though no third party has yet replicated the within-task shuffle control itself.
2026-10-05T23:34:20Z
evidence attached: reddit.post.1wyjp7i — Pre-registered experiments showing struggle-triggered escalation (+7 successes at ~1.3x cost, fact-reporting beating advice) provide independent evidence that behavioral mechanisms, not learned difficulty routing, deliver cheap-vs-expert savings.
2026-09-28T22:04:54Z
origin walked (opencode/cheap-glm, conf 0.95): anchor reddit.post.1wr7evz -> echo.blog.9a41465f6d by Brian Feeny
2026-09-28T22:02:50Z
grounded: converges/high — Converges with the Scout–Senior Split's explicit demotion of 'generic small-model/large-model routing' as an adjacent-not-sufficient pattern and the Model Barbe
2026-09-28T21:54:22Z
case created — A confound-eliminated negative result on RouteLLM's released data is a specific, resolvable mechanism claim underlying several inference-economics cases, despite near-zero engagement.
Decision trace
- 10-06 13:10repriceThe thesis moved from single-author negative result to independently corroborated: KangarooAnxious9394's pre-registered Haiku/Sonnet repo-commit experiments found struggle-triggered escalation de
- 10-06 13:10groundIndependent outsiders arrive, via a confound-eliminated negative result plus a methodologically unrelated pre-registered corroboration, exactly where Scott's canon already stands: the Scout-and-t
- 10-06 10:34attachPre-registered experiments showing struggle-triggered escalation (+7 successes at ~1.3x cost, fact-reporting beating advice) provide independent evidence that behavioral mechanisms, not learned diffic
- 10-06 10:28propose_attachPre-registered experiments showing struggle-triggered escalation (+7 successes at ~1.3x cost, fact-reporting beating advice) provide independent evidence that behavioral mechanisms, not learned diffic
- 09-29 08:04promote_anchororigin walk conf 0.95
- 09-29 08:02groundConverges with the Scout–Senior Split's explicit demotion of 'generic small-model/large-model routing' as an adjacent-not-sufficient pattern and the Model Barbell's no-middle-tier
- 09-29 07:54createA confound-eliminated negative result on RouteLLM's released data is a specific, resolvable mechanism claim underlying several inference-economics cases, despite near-zero engagement.