Reddit evaluator s1lverkin reports two replicated 100-slot Terminal-Bench 2.1 runs on public Harbor job data in which GPT-6 Luna substantially underperforms GPT-5.6 Luna on coding tasks; broad replication or OpenAI acknowledgment would establish a real coding regression contradicting GPT-6 Luna's release claims.
state: corroboratedheat: lowuncertainty: mediumconvergesscott: highindependent-model-evaluation coding-agents terminal-bench model-regressionOpenAIs1lverkin
What is this?
OpenAI released GPT-6 Sol and GPT-6 Luna on 2026-09-22, nineteen days after flagship GPT-6 Astra, positioned as halving per-task cost while keeping capability at GPT-5.6 levels. Independent evaluator Artificial Analysis partially corroborates the regression claim in the case: GPT-6 Luna's Coding Agent Index dropped 2 points vs GPT-5.6 Luna (while Sol gained 2), with mixed per-benchmark movement and notable knowledge-work regressions (GDPval-AA ~75 Elo, AA-Briefcase ~45 Elo); The Decoder summarized the release as halving prices while 'barely moving the needle.' The case's specific artifact โ s1lverkin's two replicated 100-slot Terminal-Bench 2.1 runs on public Harbor job data โ is not itself present in the supplied snippets, so its magnitude ('substantially underperforms') is uncorroborated here; vendor-published data does put GPT-5.6 Luna at 84.3% on TB 2.1, giving a baseline to regress from, and multiple Reddit threads (r/OpenAI, r/ChatGPT) independently report Luna 6 feeling like a downgrade, though one r/codex commenter reads it as parity-at-half-cost and another stresses harness-over-model effects.
Why it matters to Scott
Independently arrives at the world his evidence-class ladder and benchmark-integrity canon argue for: a vendor release claim ('half cost, GPT-5.6 capability') being adjudicated by same-harness replication on public artifacts, with Artificial Analysis already showing a โ2 coding-index move โ and the attribution question (weights regression vs harness/decoding change) is exactly his model-plus-harness benchmark unit. It also bears on what he builds: a confirmed budget-tier coding regression forces re-evaluation of his LiteLLM cheap-tier routing and barbell configs before any Luna re-pin (model-perishability discipline in action), and the bounded, resolvable artifact trail is a dated-receipts publishing probe.
ip:concept.evidence-class-ladderip:concept.model-plus-harness-benchmark-unitip:concept.model-barbellip:concept.model-perishabilitydev:technology.litellmdev:concept.cost-tiered-llm-routingdev:concept.trace-backed-agent-comparisonradar:concept.benchmark-integrityradar:concept.agent-harnessesradar:concept.model-routingradar:openai-astra-silent-rerouting-claimradar:codex-july-looping-regressionradar:notch-luna-harness-cost-migrationradar:deepseek-v4-flash-terminal-bench-replication
queries asked of Scott's wikis
- terminal-bench or coding-agent eval runs in my own projects โ which models scored what
- which models my agents pin or route to โ cheap tier / Luna / GPT-6 in production configs
- benchmark gaming and vendor release-claims credibility โ positions on eval integrity
- model regression / silent-update detection in agent pipelines
- harness vs model attribution โ how much coding-agent variance is the harness
- cheap-tier economics โ is the budget tier good enough for agentic coding
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 381h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p69 vs 1032 stories at the 336h mark (now 381h old) โ ahead of nari-qwen3-speech-ga (1.0x), behind google-cc-household-agent (1.0x)
Evidence (2) โ โญ canonical anchor
Interpretation history
2026-09-26T15:39:35Z
ElectronicAd4565's second thread adds anecdotal breadth โ Luna 6 unusable where 5.6 worked, plus scaledev's personal tests showing Sol 6 also underperforming its predecessor (in tension with AA's +2 for Sol, hinting the problem may span the whole GPT-6 cheap tier) โ but no new evaluative line, so the case's meaning is unchanged while its attention has clearly peaked: ~1.5 pts/h vs an 18.7 peak, cooling, still single-platform. Heat cools to low with the corroboration standing.
2026-09-26T15:24:09Z
evidence attached: reddit.post.1wqs9vz โ Independent user reports GPT-6 Luna unusable for coding while 5.6-Luna works unchanged, real-world echo of the replicated Terminal-Bench regression.
2026-09-26T08:32:04Z
Corroborated on substance, not engagement: s1lverkin's two same-harness runs now align with Artificial Analysis's independent โ2 coding-index move for Luna 6 vs 5.6, and the new thread comments add independent failure anecdotes in adjacent tasks (handwriting analysis, mid-project 5.6โ6.0 swap, token-efficiency vs Astra). What the hypothesis still needs is absent โ third-party same-harness replication of the magnitude, attribution (weights regression vs the harness/effort-config confound one commenter raises), any OpenAI response โ so belief moves one notch while attention moves faster: top-quintile peer velocity (13x baseline, 86th percentile) but single-platform and steady, pricing heat medium.
2026-09-25T19:34:43Z
grounded: converges/high โ Independently arrives at the world his evidence-class ladder and benchmark-integrity canon argue for: a vendor release claim ('half cost, GPT-5.6 capability') b
2026-09-25T19:26:16Z
case created โ A replicated independent evaluation with public benchmark artifacts alleging a frontier coding-model regression is a bounded, resolvable episode distinct from OpenAI's own release-claims case and from the unrelated silent-rerouting allegation.
Decision trace
- 09-27 07:40review_screenjev screen: no material development (noul=0.08)
- 09-27 07:20sensor_dirtycomment_update
- 09-27 01:39repriceElectronicAd4565's second thread adds anecdotal breadth โ Luna 6 unusable where 5.6 worked, plus scaledev's personal tests showing Sol 6 also underperforming its predecessor (in tension with
- 09-27 01:24attachIndependent user reports GPT-6 Luna unusable for coding while 5.6-Luna works unchanged, real-world echo of the replicated Terminal-Bench regression.
- 09-27 01:22propose_attachIndependent user reports GPT-6 Luna unusable for coding while 5.6-Luna works unchanged, real-world echo of the replicated Terminal-Bench regression.
- 09-26 22:41review_screenjev screen: no material development (noul=0.23)
- 09-26 20:20sensor_dirtycomment_update
- 09-26 18:32repriceCorroborated on substance, not engagement: s1lverkin's two same-harness runs now align with Artificial Analysis's independent โ2 coding-index move for Luna 6 vs 5.6, and the new thread comme
- 09-26 14:20sensor_dirtycomment_update
- 09-26 10:21sensor_dirtyvelocity_spike
- 09-26 08:22sensor_dirtycomment_update
- 09-26 05:34groundIndependently arrives at the world his evidence-class ladder and benchmark-integrity canon argue for: a vendor release claim ('half cost, GPT-5.6 capability') being adjudicated by same-harne
- 09-26 05:26createA replicated independent evaluation with public benchmark artifacts alleging a frontier coding-model regression is a bounded, resolvable episode distinct from OpenAI's own release-claims case and