2026-10-11 18:03 UTC

Independent replication will determine whether four-model orchestration in Claude Code consistently underperforms simpler single-model setups on Terminal-Bench because coordination and refusal failures outweigh specialization gains.

state: expiredheat: lowuncertainty: highconvergesscott: mediummulti-model-orchestration coding-agents agent-harnesses terminal-benchQuesmaAnthropic

What is this?

The case concerns a reported Quesma experiment wiring four models into Claude Code and testing the setup on Terminal-Bench, with the claim that coordination and refusal failures outweighed gains from model specialization. The supplied snippets establish Terminal-Bench as an agent benchmark and independently describe coordination as a central multi-agent difficulty; they also show that structured, DAG-based skill orchestration can outperform flat invocation. However, none of the snippets directly documents Quesma’s experiment or an independent replication, so the claimed consistent underperformance and its four specific failure modes remain unverified here.

Why it matters to Scott

The reported backfire converges with Scott’s view that multi-agent gains depend on explicit decomposition, routing, supervision, and cheap handoffs—not merely wiring several models together—and with his model-plus-harness benchmark unit. It directly invites a trace-backed replication relevant to his multi-model systems, but the Quesma result remains unverified and adjacent radar cases already track coordination overhead and model-routing regressions.
ip:framework.micro-agents-architectureip:concept.model-plus-harness-benchmark-unitdev:concept.deterministic-agent-control-planedev:concept.trace-backed-agent-comparisondev:concept.task-aware-model-routingradar:open-ended-agent-coordination-benchmarkradar:multi-model-orchestrator-worker-agentsradar:claude-subagent-roster-overheadradar:concept.coding-agent-benchmarks
queries asked of Scott's wikis
  • multi-agent coding versus single-agent performance
  • coordination tax in coding-agent harnesses
  • specialist agents and model-routing failure modes
  • refusal propagation in agent orchestration
  • Terminal-Bench harness design and evaluation
  • DAG-based versus flat agent orchestration

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnI wired 4 models together in Claude Code. It backfired 4 ways on Terminal-Benchbkotrys71
🟧 echo.blog ⭐Reports that connecting four models inside Claude Code backfired in four ways during Terminal-Bench testing.Quesma——
🟠 redditOpus 5 solved security tasks when I asked directly, but refused the same tasks when my orchestrator delegated them (445 Terminal-Bench trials)
ClaudeAI
Bartaseth12
🟧 hnClaude Opus 4.8 Costs 57.1× More, Loses All 5 Benchmarks – Beaten by the HarnessFloatboat_ai10
🟠 redditClaude Code Orchestrator on Terminal-Bench: Same model, same tasks - Opus refused only when the work was delegated
artificial
Bartaseth138

Interpretation history

Decision trace