2026-10-11 17:16 UTC

Official pricing and production usage will determine whether OpenAI’s reported API price cut of more than 20% for GPT-5.6 Sol materially shifts model selection or deferrable inference workloads toward its API.

state: resolvedheat: highuncertainty: lowknownscott: mediumopenai-api-pricing inference-economics frontier-modelsOpenAI

What is this?

The supplied reports describe OpenAI cutting GPT-5.6 Sol API and eligible credit pricing on August 21, 2026; Reuters cites standard short-context rates falling from $5 to $4 per million input tokens and from $30 to $20 per million output tokens, reductions of 20% and roughly 33%, respectively. Enterprise DNA reports the promotion lasts through at least November 21, while citybiz instead calculates a $24 output rate; the supplied OpenAI announcement snippet is truncated, and its July 30 blog documents an earlier change that left Sol pricing unchanged. September reports concern a separate GPT-6 release, not this price cut. The snippets establish reported price changes but provide no production-usage evidence that developers switched models or moved deferrable workloads to OpenAI.

Why it matters to Scott

Scott already holds the relevant position in Model Perishability: model repricing requires swappable providers and re-evaluation; this reported cut gives a concrete reason to recheck his LiteLLM routing economics and paid OpenAI usage, rather than establishing a new strategic claim. Conflicting price reports and absent production evidence prevent concluding that Sol should win more workloads; the supplied radar pages track related inference economics, not this exact cut.
ip:concept.model-perishabilityip:concept.ai-unit-economicsdev:concept.task-aware-model-routingdev:technology.litellmwork:project.openairadar:concept.inference-economicsradar:concept.model-routingradar:deepseek-v4-time-of-day-pricing
queries asked of Scott's wikis
  • model routing cost per successful task evaluation
  • coding agent harness token budgets provider selection
  • deferrable inference batch workloads scheduling economics
  • local open models versus hosted API economics
  • inference price declines demand elasticity Jevons paradox

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

08-20 14:00⭐ origin echo-reconstructedOpenAI posted: “As we continue to push the frontier of capabilities while improving efficiency, we're dropping API and credit pricing of GPT
OpenAI on x (echo) · attributed from reddit.post.1vuxlw7
—
08-22 00:33first on r/OpenAI · published · +34.6hAPI Price cut > 20%
Babayaga1664
—
08-22 04:33first on hacker news · published · +38.5hGPT 5.6 Sol 20% price reduction
izakfr
—
09-10 18:08first on r/singularity · published · +508.1hGPT-6 Sol Appeared on the OpenAI API
141_1337
—
09-22 18:00first on openai · published · +796.0hIntroducing GPT-6 Sol and Luna
OpenAI
—
09-23 07:34first on r/ClaudeAI · published · +809.6hClaude Opus 5.5 can’t really be considered a cheaper model than OpenAI new Sol model. It’s roughly twice as expensive as GPT‑6 Sol, yet it only reduces output tokens by about 19% per task and cuts reasoning/indexing tokens by about half.
RFOK
—
08-22 00:33amplified on r/OpenAIreddit.post.1vuxlw7
Babayaga1664
peak 6 · 2 comments · 0% of case engagement
08-22 04:33amplified on hacker newshn.story.49396590
izakfr
peak 90 · 77 comments · 8% of case engagement
08-24 15:14amplified on r/OpenAIreddit.post.1vx5mrz
Dualyeti
peak 175 · 68 comments · 6% of case engagement
08-24 15:22amplified on hacker newshn.story.49421074
tosh
peak 338 · 322 comments · 32% of case engagement
08-26 00:11amplified on hacker newshn.story.49442514
Airealist
peak 1 · 0 comments · 0% of case engagement
08-28 15:04amplified on hacker newshn.story.49479633
Philpax
peak 1 · 0 comments · 0% of case engagement
8 more amplifiers in ainews.case_chain
08-22 01:20our radar first saw it · +35.3hdiscovery anchor: reddit.post.1vuxlw7—

Evidence (16) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditAPI Price cut > 20%
OpenAI
Babayaga166462
🟧 echo.x ⭐OpenAI posted: “As we continue to push the frontier of capabilities while improving efficiency, we're dropping API and credit pricing of GPTOpenAI——
🟧 hnGPT 5.6 Sol 20% price reductionizakfr9074
🟠 redditNew Sol API pricing - $4 per million input tokens and $20 per million output tokens
OpenAI
Dualyeti17568
🟧 hnOpenAI: GPT 5.6 Sol price reduction (until at least Nov 21)tosh338319
🟧 hnAI Realist Radar:GPT‑5.6 Sol Pricing, Stripe's OpenRouter DealAirealist10
🟧 hnGPT 5.6 Discounts and Jevons ParadoxPhilpax10
🟠 redditGPT-6 Sol Appeared on the OpenAI API
singularity
141_133745090
🟧 hnGPT-6-sol appeared on OpenAI APIpranshuchittora107
🟠 redditIntroducing GPT-6 Sol and Luna
OpenAI
DemiPixel1128189
🟧 openaiIntroducing GPT-6 Sol and Luna
Retrieved article excerpt

Open article · Retrieved 2026-09-22T18:24:17.863121+00:00

# Introducing GPT‑6 Sol and Luna

More ways to bring frontier intelligence into the work you do every day.

Loading…

Share

Earlier this month, we introduced [GPT‑6 Astra](https://openai.com/index/gpt-6-astra/), the most intelligent and aligned model in the world. While the most demanding and important projects still call for Astra’s full depth, work happens at different scales, rhythms, and budgets.

That’s why we’re expanding the GPT‑6 universe with **GPT‑6 Sol** and **GPT‑6 Luna.** GPT‑6 Astra introduced a new generation of intelligence—these models help distribute the benefits of that intelligence by advancing the frontier on cost efficiency. We trained GPT‑6 Sol and Luna with similar methods as GPT‑6 Astra, bringing the advances behind Astra’s state-of-the-art performance in professional work, factuality, coding, computer use, and alignment to faster, more affordable models.

The GPT‑6 models lead across the cost–intelligence curve, combining exceptional capabilities at every tier with infrastructure that delivers them efficiently at scale. Improvements in caching and inference let us serve these models at lower cost, and we’re passing those savings directly on to users and customers by **reducing API prices for Sol and Luna by 50%** compared with their GPT‑5.6 promotional pricing. Together, these improvements make advanced AI practical for more everyday tasks and applications at scale.

### GPT‑6 API pricing

|  |  |  |  |
| --- | --- | --- | --- |
| **Model** | **Input** | **Output** | **Price reduction** |
| **GPT‑6 Sol** vs. GPT‑5.6 Sol | $4 → **$2** | $20 → **$10** | 50% cheaper |
| **GPT‑6 Luna** vs. GPT‑5.6 Luna | $0.20 → **$0.10** | $1.20 → **$0.50** | 50% cheaper |

*Prices are per 1 million tokens.*

**GPT‑6 Astra** continues to be our best model across the board. Choose it when you want the best results and an uncompromising experience.

## A step up across the model family

GPT‑6 Sol and Luna bring intelligence upgrades and cost efficiency to the models you already know and use across capabilities most useful for getting complex work done.

### Professional work

GPT‑6 Sol can take on difficult work tasks while giving you more room to iterate with higher usage limits and lower cost, offering more intelligence and better results versus similarly priced competitor models.

On **AutomationBench,** a test of business workflows across apps, GPT‑6 Sol at xhigh effort outperforms Claude Opus 5 at max effort at just 9% of Opus 5’s cost per task. At high effort, GPT‑6 Luna improves on its predecessor by 5.4 percentage points at 58% lower cost per task.

*In* [*AutomationBench 1.0.6*⁠(opens in a new window)](https://zapier.com/benchmarks)*, AI agents are tested on end-to-end workflows using 47 tools across sales, marketing, operations, support, finance, and HR. The datapoint for Claude Fable 5.1 understates its actual cost, as it omits the cost of the Opus 5 fallbacks, which occurred on ~40% of tasks.*

GPT‑6 Sol also exceeds Claude Fable 5.1 at far lower cost, and even bests low-effort GPT‑6 Astra.

| **Model (and effort)** | **Score** | **Cost per task** |
| --- | --- | --- |
| GPT‑6 Sol (xhigh) | 33.2% | $0.27 |
| GPT‑6 Astra (low) | 30.3% | **3.9x** GPT‑6 Sol |
| Claude Opus 5 (max) | 26.9% | **11.1x** GPT‑6 Sol |
| Claude Fable 5.1 w/ Opus 5 Fallback (max) | 31.4% | **>8.9x** GPT‑6 Sol  *(fallback cost not reported)* |

On **Agents’ Last Exam**, which evaluates agents on complex professional workflows, GPT‑6 Sol at max effort scores 56.4%, above Claude Opus 5’s highest score in the evaluation at 60% lower cost per task.

*In* [*Agents’ Last Exam V1*⁠(opens in a new window)](https://agents-last-exam.org/)*, AI agents are evaluated on long-horizon, economically valuable tasks spanning 55 sub-industries, covering most major fields of professional work performed on a computer.*

### Factuality

The usefulness of an answer depends on getting the facts right, and we’re continuing to make progress on factual reliability. On our internal factuality evaluation, which is based on de-identified real-world conversations where users flagged mistakes by our models, GPT‑6 Sol makes about half as many mistakes as its predecessor, approaching Astra-level reliability at much lower cost. GPT‑6 Luna also improves substantially; at higher effort levels it matches GPT‑5.6 Sol at about a hundredth its cost.

*Here we evaluate factuality on de-identified ChatGPT conversations where users had flagged a factual error from a prior model. These error-inducing conversations are not representative of typical usage, where factual errors are more rare. Scores are not controlled for length; however, our verbosity sweeps showed almost no dependence on answer length.*

### Coding

This year, coding agents have begun tackling tasks with more complexity, scope, and duration than ever before. At OpenAI, our internal usage has grown exponentially. Valued at API prices, daily token usage has exceeded $600 for the median researcher and $7,000 for researchers at the 90th percentile ([Research acceleration: The view inside OpenAI⁠](https://openai.com/index/research-acceleration-view-inside-openai/)). As coding agents take on longer and more demanding tasks, the cost of sustained use matters more. GPT‑6 Sol and Luna combine strong coding performance with lower API prices, giving developers more room to iterate and teams the confidence to be more ambitious about what they ask Codex to take on.

On **FrontierCode**, which evaluates whether coding agents produce changes ready to merge into real codebases, GPT‑6 Sol improves substantially over GPT‑5.6 Sol, and is able to match Claude Fable 5.1 xhigh at much lower cost.

*In* [*FrontierCode 1.1 Main*⁠(opens in a new window)](https://cognition.com/frontiercode)*, AI agents write code that’s graded not only on correctness but also “mergeability”: e.g., test quality, scope discipline, code style, and adherence to codebase standards.*

On **DeepSWE v1.1,** which tests performance on complex software-engineering tasks in real codebases, GPT‑6 Sol at max effort scores 68.8%, within 1.1 percentage points of Claude Fable 5’s highest score in the evaluation—69.9% at xhigh effort—at approximately 80% lower cost per task.

GPT‑6 Luna at max effort scores 66.6%, comparable to Claude Opus 5 and Fable 5 at medium effort. In these comparisons, Luna costs 93% less per task than Opus 5 and 96% less than Fable 5.

*In* [*DeepSWE 1.1*⁠(opens in a new window)](https://deepswe.datacurve.ai/)*, AI agents solve original, long-horizon software engineering tasks.*

### Computer use

While GPT‑6 Astra remains the world’s best model for computer use, GPT‑6 Sol and Luna offer more cost-efficient performance than their predecessors. On **OSWorld 2.0 offline**, GPT‑6 Sol at xhigh effort achieves a similar score to Claude Opus 5 at medium effort—60.5% versus 60.3%—at approximately 80% lower cost per task. GPT‑6 Luna (max) is able to exceed GPT‑5.6 Sol (medium) at one tenth of its cost.

*In* [*OSWorld 2.0*⁠(opens in a new window)](https://osworld-v2.xlang.ai/)*, AI agents attempt long-horizon computer-use workflows spanning everyday and professional tasks. We report the partial reward on the offline set from the v2026.08.08 release.*

### Collaboration style

We’ve also brought GPT‑6 Astra’s improved communication style to Sol and Luna, which we think will be especially noticeable in technical and coding conversations. Expect to see more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall without losing substance.

Prompt

Website’s looking clean! Could we use a Bento Box design style and add a slider between the pages in the top right? You may need to dive into the React..

GPT-5.6 Sol

Absolutely. I can get the bento feel and the top-right sliding page switcher without pulling in React—the interaction is small enough to stay fast and smooth in the current site. I’m reshaping all four pages around that system now.

---

Done — open the live site.

It now has a colorful bento-box layout across all four pages and an animated top-right Home / Play / Place / Taste slider. React wasn’t necessary; native page transitions keep it lightweight and smooth.

The built-in image tool created the food artwork. Prompt: “An energetic gouache-and-ink editorial illustration celebrating dumplings and spicy noodles, with bright colors, steam, spice, and no text or logos.”

GPT-6 Sol

I’ll reshape the pages into a bento grid and turn the top-right links into a sliding page switcher. I’ll keep the existing site and check whether that interaction needs React before changing its setup.

---

The updated site is live. Every page now has a playful bento layout, and the top-right page switcher slides between sections. I checked it on desktop and narrow mobile screens, including browser back navigation. The existing site didn’t need React for this.

*Although style is subjective, we prefer GPT‑6 Sol’s reply here. It doesn’t jump to conclusions as quickly, spends less time reiterating details that might be obvious to the asker (e.g., that the website has four pages), uses less vague language (e.g., “bento feel”, “reshaping… around that system”), is more forthcoming with what it did and didn’t check, and doesn’t unnecessarily share implementation details like its image tool prompt.*

## Improving caching for agents and long conversations

Alongside lower token prices, we’re helping developers building on GPT‑6 save more on the context their applications reuse. We’ve improved prompt caching for GPT‑6 to deliver higher cache hit rates by default, helping agents reuse more context, respond faster, and benefit from discounts of 90% on cached input-token reads.

Developers also have more ways to measure and optimize their caching performance:

- **Monitor and diagnose.** The [Prompt Caching Dashboard⁠(opens in a new window)](https://platform.openai.com/usage?usage_section=prompt-caching) shows how much input is cached and how that changes over time. The [diagnostics tool⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics) helps explain missed opportunities for caching and what to fix.
- **Adjust reasoning effort and tool availability without breaking cache.** Increase [reasoning effort⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/reasoning#change-reasoning-mid-conversation) for harder tasks or lower it for simpler follow-ups, and [enable or disable tools⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching#how-to-optimize-prompt-caching) as your agent’s needs change. Both controls now preserve earlier context for cache reuse.
- **Optimize which prefixes get cached.** Explicit breakpoints let developers choose where cached prompt prefixes end. This gives developers more control over cache reuse and can improve performance.

GitHub reports that, over the past several months, these improvements have reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests to OpenAI models, helping Copilot respond faster.

## Continuing to improve alignment

GPT‑6 Sol and Luna build on the alignment work introduced with Astra, our most aligned model to date. In our alignment evaluations, both Sol and Luna show improvements over their GPT‑5.6 counterparts, including lower rates of misleading claims about their coding work.

The evaluations below deliberately test challenging situations and do not measure failure rates in typical use. See the [system card⁠(opens in a new window)](https://deploymentsafety.openai.com/gpt-6-astra) for the full results.

## Availability

GPT‑6 Sol and GPT‑6 Luna are available in ChatGPT Work and Codex starting today for all Plus, Pro, Business, Enterprise, and Edu users. Free and Go users can access GPT‑6 Luna in the desktop app. These models are not yet available in Chat. In the OpenAI API, they are 
OpenAI——
🟧 hnGPT-6 Sol and Luna push the cost-efficiency frontierwertyk40
🟠 redditTheory: Sol is the new Terra and should be compared to Sonnet not Opus
OpenAI
TraditionalHome88526133
🟧 hnGPT-6 Sol and Luna push the cost efficiency frontier by halving token costtheanonymousone10
🟠 redditClaude Opus 5.5 can’t really be considered a cheaper model than OpenAI new Sol model. It’s roughly twice as expensive as GPT‑6 Sol, yet it only reduces output tokens by about 19% per task and cuts reasoning/indexing tokens by about half.
ClaudeAI
RFOK07
🟠 redditCost Efficiency chart of the GPT-6 Models
OpenAI
Tall_Abrocoma_35333910

Interpretation history

Decision trace