The case concerns a reported Quesma experiment wiring four models into Claude Code and testing the setup on Terminal-Bench, with the claim that coordination and refusal failures outweighed gains from model specialization. The supplied snippets establish Terminal-Bench as an agent benchmark and independently describe coordination as a central multi-agent difficulty; they also show that structured, DAG-based skill orchestration can outperform flat invocation. However, none of the snippets directly documents Quesma’s experiment or an independent replication, so the claimed consistent underperformance and its four specific failure modes remain unverified here.
The reported backfire converges with Scott’s view that multi-agent gains depend on explicit decomposition, routing, supervision, and cheap handoffs—not merely wiring several models together—and with his model-plus-harness benchmark unit. It directly invites a trace-backed replication relevant to his multi-model systems, but the Quesma result remains unverified and adjacent radar cases already track coordination overhead and model-routing regressions.
ip:framework.micro-agents-architectureip:concept.model-plus-harness-benchmark-unitdev:concept.deterministic-agent-control-planedev:concept.trace-backed-agent-comparisondev:concept.task-aware-model-routingradar:open-ended-agent-coordination-benchmarkradar:multi-model-orchestrator-worker-agentsradar:claude-subagent-roster-overheadradar:concept.coding-agent-benchmarks
queries asked of Scott's wikis
- multi-agent coding versus single-agent performance
- coordination tax in coding-agent harnesses
- specialist agents and model-routing failure modes
- refusal propagation in agent orchestration
- Terminal-Bench harness design and evaluation
- DAG-based versus flat agent orchestration
2026-08-14T11:29:42Z
Repeated reobservation has produced no independent run, controlled baseline, traces, or root-cause artifact; the single-operator result remains uncorroborated and the episode has faded pending genuinely new replication.
2026-08-12T10:31:20Z
The refreshed comments reinforce that delegation changes prompts, permissions, context, and refusal surfaces, and identify artifacts a valid replication should log. They add no independent benchmark run or controlled comparison, so the core underperformance hypothesis remains uncorroborated.
2026-08-12T08:37:08Z
The newly attached Reddit post is a restatement by the same operator of the same 445-trial experiment, not an independent replication. It sharpens the delegated-context refusal observation but leaves the broader claim of consistent multi-model underperformance uncorroborated.
2026-08-12T08:22:33Z
evidence attached: reddit.post.1vm7a6t — This independent Terminal-Bench report directly bears on whether delegated multi-model orchestration causes refusals and underperformance.
2026-08-11T11:44:30Z
The new link is only an adjacent claim that harness design can outweigh model choice; without methodology, results, or evidence of an independent Terminal-Bench replication, it does not corroborate the four-model backfire hypothesis.
2026-08-11T11:23:06Z
evidence attached: hn.story.49256231 — The report adds evidence that harness design and orchestration can dominate model choice in coding-agent benchmarks.
2026-08-10T22:41:37Z
The Reddit attachment turns the bare claim into a concrete single-operator report covering 445 Terminal-Bench trials, with benchmark conditions and model roles described. It remains the same experiment rather than an independent replication, and without traces or a controlled single-model comparison it does not establish consistent orchestration underperformance.
2026-08-10T22:22:55Z
evidence attached: reddit.post.1vknpd1 — shared external link with case evidence
2026-08-10T15:37:36Z
No replication, methods, traces, or benchmark scores have emerged; the unchanged observation leaves this as a single self-reported result rather than evidence of consistent multi-model underperformance.
2026-08-10T15:31:05Z
grounded: converges/medium — The reported backfire converges with Scott’s view that multi-agent gains depend on explicit decomposition, routing, supervision, and cheap handoffs—not merely w
2026-08-10T15:28:14Z
case created — The report presents a bounded negative benchmark result on a consequential coding-agent harness tradeoff that is suitable for replication.