2026-10-11 18:02 UTC

Independent use will confirm whether frontier-model orchestration with cheaper worker models preserves most coding-agent performance while cutting inference cost by roughly half.

state: resolvedheat: lowuncertainty: mediumknownscott: mediummulti-model-orchestration coding-agents agent-costsAnthropicOpenAI

What is this?

The case concerns a hybrid coding-agent architecture in which a frontier model handles planning or orchestration while cheaper models execute delegated work, aiming to retain quality while reducing inference costs. The supplied snippets support growing interest in such hybrid workflows, rising coding-agent costs, and a narrowing performance gap between frontier and cheaper or self-hostable models. However, they do not substantiate the named “Fable 5” system, Anthropic’s alleged benchmark, the specific 96%-performance/46%-cost figures, or the claim of independent confirmation.

Why it matters to Scott

Scott already holds this architecture in “Scout–Senior Split” and “Model Barbell,” and implements related task-aware routing in active agent systems. A reproducible coding benchmark could quantify or challenge those claims and influence routing economics, but the supplied material does not establish the named benchmark or its 96%-performance/46%-cost figures, so relevance remains medium pending independent confirmation.
ip:framework.scout-senior-splitip:concept.model-barbelldev:concept.task-aware-model-routingip:concept.evaluation-driven-developmentdev:project.proposalradar:concept.agent-harnessradar:concept.coding-agentsradar:concept.multi-agent-coordination
queries asked of Scott's wikis
  • frontier planner cheap worker model orchestration
  • coding-agent model routing and delegation
  • cost-aware agent harness architecture
  • coding-agent quality versus inference cost
  • local models as agent workers
  • multi-model evaluation for software tasks

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (29) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Anthropic just benchmarked "Fable 5 orchestrates, cheap models execute": 96% of the performance at 46% of the cost. You can run this pattern in Claude Code today
ClaudeAI
john9901291400236
🟠 redditFable + 5.6 is absolute peak
ClaudeCode
Bright-Celery-40581118199
🟧 hnStandalone Pi coding agent extension harness for two-model agentic engineeringcromka11
🟠 redditHermes/OpenClaw vs Codex Remote/Claude Dispatch
LocalLLaMA
stevyhacker07
🟠 redditLocal models as sub-agents with cloud orchestrators?
LocalLLaMA
neeeser014
🟠 redditI built an MCP server so Claude Code can delegate work to GPT-5.6, DeepSeek, GLM and a local Qwen — then benchmarked all of them against Claude itself (198 runs, hidden tests)
ClaudeAI
MeetStraight189917898
🟧 hnAgent swarms and the new model economicsjlaneve257127
🟧 hnRamp opens AI model router, says it cut internal LLM costs 30%ryanmerket10
🟧 hnYou only need the frontier model for one single editjxmorris1222393
🟠 redditWorkflow: I'm letting Opus decide when and how to use Fable
ClaudeAI
Odd_Sale381111
🟠 redditAgent swarms and the new model economics
singularity
petburiraja171
🟧 hnOpen-ultra: a self-training LLM routing proxyjoshkolo20
🟧 hnShow HN: Adversarial code review setup with herdr, Claude and GPT-5.6-soloverflowy21
🟧 hnAI Coding Assistant Strategy Slashes Token Billsrbanffy10
🟠 redditRound 3: the comment section designed my benchmark — 13 lanes, controlled reasoning effort, and a knowledge-cutoff trap. The cheap models didn't fail at reasoning; they failed at knowing what year it is.
ClaudeAI
MeetStraight189974
🟧 hnReducing LLM Costs 50% Using Best-Execution for Intelligencearbayi10
🟠 redditI built Frugal: a plugin that routes Claude Code work to the cheapest model that can do it
ClaudeAI
StolenDrinks5325
🟠 redditUsing Claude Code and Antigravity CLI together. Having issues
ClaudeAI
CommandCheap795018
🟧 hnBuilding an AI-orchestrated publishing workflow for a long-form writing projecttmuhlestein10
🟠 redditCactus Hybrid: We taught Gemma 4 to know when it's wrong
LocalLLaMA
Henrie_the_dreamer18243
🟧 hnShow HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrongHenryNdubuaku18340
🟧 hnAI-maestro: Conduct a roster of AI coding agents against a work boardmychiefmind10
🟧 hnShow HN: Echo – Fable-level results at 1/3 the cost using open-weight modelsadam_rida20
🟠 redditI built a router that spreads work between my Claude and ChatGPT subscriptions
ClaudeAI
YaBoyChips381912
🟧 hnThe Subagent Taxsystima10
🟧 hnShow HN: Echo – Fable-level results at 1/3 the cost using open-weight modelsadam_rida444213
🟧 hnShow HN: Run Claude Code Through Codex, Kimi, Grok, or Cursorrane10
🟠 redditWhat exactly are the benefits of using agents? Because I have outright banned it.
ClaudeAI
MysteriousFloor1406119
🟠 redditWe spend $0.001 to decide if we need to spend $0.08
ClaudeAI
Dizzy_Leg_591212

Interpretation history

Decision trace