2026-10-11 16:37 UTC

bfeeny's controlled experiment claims a learned cheap-vs-expert router scoring 0.84 held-out AUC still scores 0.838 when within-task labels are shuffled — it learned task identity, not difficulty — leaving output-based deferral, not learned routing, as the practical cost-saving mechanism until replication finds genuine difficulty signal.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highllm-routing inference-economics evaluation-methodology

What is this?

This case tracks a negative-result replication claim by Brian Feeny (bfeeny): his 21 Sep 2026 post 'Replicating RouteLLM on Amazon Bedrock' reports that a learned cheap-vs-expert router with 0.84 held-out AUC still scores ~0.838 when cheap/expert labels are shuffled within each task — i.e. the router learned which task (hence which model) produced a query, not per-query difficulty — so the practical cost-saving mechanism is output-based deferral (run the cheap model, escalate on failure) rather than learned routing. A second, methodologically unrelated line (KangarooAnxious9394, Oct 2026, pre-registered Haiku-cheap/Sonnet-strong experiments on real repo commits) independently corroborates the behavioral half: struggle-triggered escalation netted +7 successes at ~1.3x cost, with fact-reporting beating advice. Strong caveats: the supplied web results contain nothing on this work (they return unrelated MoE-router and generic-ML papers), the within-task shuffle control is single-author and unreplicated by third parties, both posts drew near-zero community scrutiny (score 0, upvote ratios 0.17–0.2), and the canonical blog anchor is an echo-reconstructed testimony source — the numbers rest entirely on the authors' own reporting.

Why it matters to Scott

Independent outsiders arrive, via a confound-eliminated negative result plus a methodologically unrelated pre-registered corroboration, exactly where Scott's canon already stands: the Scout-and-the-Senior demotion of generic small/large-model routing and his default-cheap-with-fallback designs (cost-tiered routing, metacognitive failure escalation, the propose–finalise gate) all bet on behavioral deferral over learned difficulty — this is the dated receipt for that bet. The within-task shuffle control is itself the reusable contribution: a cheap falsification test to demand before crediting any learned/adaptive router's claimed savings, though it remains single-author and unreplicated, so it arms his skepticism rather than settles the question.
dev:concept.cost-tiered-llm-routingdev:concept.task-aware-model-routingdev:concept.metacognitive-resolution-controldev:concept.propose-finalise-gateradar:ramp-thompson-sampling-model-routingradar:dynamic-model-switching-evaluationradar:world-model-optimizer-agent-routingradar:nemo-switchyard-llm-routerradar:yhahn-agent-escalation-ergonomics
queries asked of Scott's wikis
  • scout-senior split small-model large-model routing demotion
  • model barbell default-cheap strong-model fallback design
  • agent escalation on failure harness struggle detection
  • evaluation confounds shortcut learning benchmark leakage
  • inference economics cost-saving claims skepticism
  • router difficulty signal vs output-based deferral

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p50momentum: steady2 platformsage 506h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-20 14:00⭐ origin echo-reconstructedBlog post "Replicating RouteLLM on Amazon Bedrock" (21 Sep 2026): "In-distribution AUC on the cascade labels looks healthy. It is 0.744 on t
Brian Feeny on blog (echo) · attributed from reddit.post.1wr7evz
—
09-27 01:35first on r/MachineLearning · published · +155.6hA learned LLM router scored 0.84 AUC. Shuffling the labels within each task still scored 0.838 [R]
bfeeny
—
10-05 20:42first on r/ClaudeAI · published · +366.7hI tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't
KangarooAnxious9394
—
09-27 01:35amplified on r/MachineLearning 👑reddit.post.1wr7evz
bfeeny
peak 0 · 1 comments · 55% of case engagement
10-05 20:42amplified on r/ClaudeAIreddit.post.1wyjp7i
KangarooAnxious9394
peak 0 · 1 comments · 55% of case engagement
09-28 21:20our radar first saw it · +199.3hdiscovery anchor: reddit.post.1wr7evz—
pace: p9 vs 1032 stories at the 336h mark (now 506h old) — behind addom-local-coding-harness (0.5x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditA learned LLM router scored 0.84 AUC. Shuffling the labels within each task still scored 0.838 [R]
MachineLearning
Retrieved article excerpt

Open article · Retrieved 2026-09-28T21:36:38.938201+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
bfeeny01
🟧 echo.blog ⭐Blog post "Replicating RouteLLM on Amazon Bedrock" (21 Sep 2026): "In-distribution AUC on the cascade labels looks healthy. It is 0.744 on tBrian Feeny——
🟠 redditI tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't
ClaudeAI
KangarooAnxious939401

Interpretation history

Decision trace