2026-10-11 16:38 UTC

Reddit evaluator s1lverkin reports two replicated 100-slot Terminal-Bench 2.1 runs on public Harbor job data in which GPT-6 Luna substantially underperforms GPT-5.6 Luna on coding tasks; broad replication or OpenAI acknowledgment would establish a real coding regression contradicting GPT-6 Luna's release claims.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highindependent-model-evaluation coding-agents terminal-bench model-regressionOpenAIs1lverkin

What is this?

OpenAI released GPT-6 Sol and GPT-6 Luna on 2026-09-22, nineteen days after flagship GPT-6 Astra, positioned as halving per-task cost while keeping capability at GPT-5.6 levels. Independent evaluator Artificial Analysis partially corroborates the regression claim in the case: GPT-6 Luna's Coding Agent Index dropped 2 points vs GPT-5.6 Luna (while Sol gained 2), with mixed per-benchmark movement and notable knowledge-work regressions (GDPval-AA ~75 Elo, AA-Briefcase ~45 Elo); The Decoder summarized the release as halving prices while 'barely moving the needle.' The case's specific artifact โ€” s1lverkin's two replicated 100-slot Terminal-Bench 2.1 runs on public Harbor job data โ€” is not itself present in the supplied snippets, so its magnitude ('substantially underperforms') is uncorroborated here; vendor-published data does put GPT-5.6 Luna at 84.3% on TB 2.1, giving a baseline to regress from, and multiple Reddit threads (r/OpenAI, r/ChatGPT) independently report Luna 6 feeling like a downgrade, though one r/codex commenter reads it as parity-at-half-cost and another stresses harness-over-model effects.

Why it matters to Scott

Independently arrives at the world his evidence-class ladder and benchmark-integrity canon argue for: a vendor release claim ('half cost, GPT-5.6 capability') being adjudicated by same-harness replication on public artifacts, with Artificial Analysis already showing a โˆ’2 coding-index move โ€” and the attribution question (weights regression vs harness/decoding change) is exactly his model-plus-harness benchmark unit. It also bears on what he builds: a confirmed budget-tier coding regression forces re-evaluation of his LiteLLM cheap-tier routing and barbell configs before any Luna re-pin (model-perishability discipline in action), and the bounded, resolvable artifact trail is a dated-receipts publishing probe.
ip:concept.evidence-class-ladderip:concept.model-plus-harness-benchmark-unitip:concept.model-barbellip:concept.model-perishabilitydev:technology.litellmdev:concept.cost-tiered-llm-routingdev:concept.trace-backed-agent-comparisonradar:concept.benchmark-integrityradar:concept.agent-harnessesradar:concept.model-routingradar:openai-astra-silent-rerouting-claimradar:codex-july-looping-regressionradar:notch-luna-harness-cost-migrationradar:deepseek-v4-flash-terminal-bench-replication
queries asked of Scott's wikis
  • terminal-bench or coding-agent eval runs in my own projects โ€” which models scored what
  • which models my agents pin or route to โ€” cheap tier / Luna / GPT-6 in production configs
  • benchmark gaming and vendor release-claims credibility โ€” positions on eval integrity
  • model regression / silent-update detection in agent pipelines
  • harness vs model attribution โ€” how much coding-agent variance is the harness
  • cheap-tier economics โ€” is the budget tier good enough for agentic coding

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 381h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-25 19:14โญ origin directly observedI ran 100 Terminal-Bench 2.1 slots on Luna 5.6 and Luna 6. The Luna 6 results still look like a joke
s1lverkin on r/OpenAI
โ€”
09-26 14:43first on r/OpenAI ยท published ยท +19.5hIt's true
ElectronicAd4565
โ€”
09-25 19:14amplified on r/OpenAI ๐Ÿ‘‘reddit.post.1wq54g2
s1lverkin
peak 60 ยท 12 comments ยท 63% of case engagement
09-26 14:43amplified on r/OpenAIreddit.post.1wqs9vz
ElectronicAd4565
peak 26 ยท 17 comments ยท 38% of case engagement
09-25 19:20our radar first saw it ยท +0.1hdiscovery anchor: reddit.post.1wq54g2โ€”
pace: p69 vs 1032 stories at the 336h mark (now 381h old) โ€” ahead of nari-qwen3-speech-ga (1.0x), behind google-cc-household-agent (1.0x)

Evidence (2) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญI ran 100 Terminal-Bench 2.1 slots on Luna 5.6 and Luna 6. The Luna 6 results still look like a joke
OpenAI
s1lverkin6012
๐ŸŸ  redditIt's true
OpenAI
ElectronicAd45652517

Interpretation history

Decision trace