2026-10-11 17:12 UTC

OpenAI claims its released GPT-6 Sol and Luna improve coding and professional-agent performance while cutting API prices roughly in half versus GPT-5.6 promotional rates, materially lowering sustained agent-work costs.

state: resolvedheat: lowuncertainty: lowconvergesscott: mediumfrontier-models coding-agents inference-economics prompt-cachingOpenAI

What is this?

OpenAI released GPT-6 Sol and Luna on 22 September 2026 — two new tiers of the GPT-6 family (flagship Astra shipped early September; Sol/Luna debuted in July as GPT-5.6 models) — with Sol aimed at complex coding, agentic and computer-control work and Luna at high-volume lightweight tasks. Pricing is half the GPT-5.6 promotional rates (Sol $2/$10 vs $4/$20; Luna $0.10/$0.50 vs $0.20/$1.20 per 1M tokens) plus a 90% cached-input discount, live day-one in ChatGPT Work, Codex and the API; Artificial Analysis independently confirms the ~50% effective cost drop ($1.06 vs $1.99 per Intelligence Index run) but notes the savings come entirely from the price cut since output tokens per task rose slightly. OpenAI's quality claims — roughly half as many factual errors (on an internal eval it acknowledges is non-representative), lower rates of coding agents misrepresenting completed work, and cost-per-task wins over Claude Opus 5 and Fable 5 — are vendor-reported only, and notably it benchmarked Opus 5 rather than the Opus 5.5 Anthropic shipped about 90 minutes before launch. The supplied web results cover only the 22 September launch: the GPT-6.1 Sol point release and its reported 'one-fifth of Astra' pricing that the case's recent evidence and reprice decision hinge on do not appear in this search and remain uncorroborated beyond Reddit testimony.

Why it matters to Scott

OpenAI's 90–95% cached-input discounts, explicit breakpoints, and near-free Luna independently arrive where Prefix-Caching Economics and the Model Barbell already sit — a vendor price sheet as dated receipt — while the halved Sol/Luna rates force a concrete reprice of his LiteLLM cheap/medium tier aliases with 6.1 Sol, not launch 6-Sol, as the target. The unresolved quality dispute (testimonial regression and fabrication-of-completed-work reports against vendor benchmarks that cherry-picked Opus 5, with 6.1 first-party pricing still unconfirmed per grounding) makes the swap a capability-audit decision — wait for an independent coding-agent eval or his own validation before treating 6.1 Sol as a 5.6 drop-in — and the week-later point release ('Terra rebranded') is fresh Model-Perishability evidence that his swap-ready gateway posture is the right one.
ip:concept.prefix-caching-economicsip:concept.model-barbellip:framework.scout-senior-splitip:concept.model-perishabilityip:concept.capability-auditdev:technology.litellmdev:concept.cost-tiered-llm-routingradar:gpt6-prompt-cache-controlsradar:openai-gpt56-sol-api-price-cutradar:anthropic-opus55-cache-read-repricingradar:coding-agent-self-report-failure-blindnessradar:concept.inference-economicsradar:concept.model-pricingradar:concept.prompt-caching
queries asked of Scott's wikis
  • model barbell scout-senior routing cheap frontier tier
  • prefix caching economics cache discount breakpoints
  • LiteLLM tier alias reprice gateway config
  • coding agent fabrication completed-work validation harness
  • frontier price collapse cost-per-task frontier shift
  • vendor benchmark baseline cherry-picking independent eval

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

09-22 18:04⭐ origin directly observedIntroducing GPT-6 Sol and Luna
DemiPixel on r/singularity
—
09-22 18:00first on hacker news · published · +-0.1hGPT-6 Sol and Luna
OfficialTurkey
—
09-22 18:04first on r/singularity · published · +0.0hGPT 6 Sol and Luna Prices
Ok_Barracuda_1161
—
09-22 19:12first on r/OpenAI · published · +1.1hOpenAI launches GPT-6 Sol + Luna: Astra intelligence moves down the cost curve
etherd0t
—
09-22 19:25first on blog (echo) · first seen by us · +1.3hAnnounces Sol and Luna availability in Work, Codex, and the API, lower token pricing, improved prompt caching, and claimed benchmark cost-pe
OpenAI
—
09-29 20:11first on r/ClaudeAI · published · +170.1hOpus 5.5 still leads on AA intelligence, but GPT-6.1 Sol shifts the cost frontier
rhiever
—
09-22 18:00amplified on hacker news 👑hn.story.49805509
OfficialTurkey
peak 1723 · 819 comments · 50% of case engagement
09-22 18:04amplified on r/singularityreddit.post.1wngzes
Ok_Barracuda_1161
peak 18 · 3 comments · 0% of case engagement
09-22 18:04amplified on r/singularityreddit.post.1wngzou
DemiPixel
peak 1334 · 313 comments · 18% of case engagement
09-22 18:06amplified on r/singularityreddit.post.1wnh1q1
ees-h
peak 28 · 10 comments · 0% of case engagement
09-22 18:15amplified on r/singularityreddit.post.1wnha3r
No_Hovercraft6239
peak 131 · 15 comments · 2% of case engagement
09-22 18:15amplified on r/singularityreddit.post.1wnhapx
141_1337
peak 31 · 8 comments · 0% of case engagement
14 more amplifiers in ainews.case_chain
09-22 18:20our radar first saw it · +0.3hdiscovery anchor: reddit.post.1wngzou—
09-23 18:15reached heat=high · +24.2h · via ledger——

Evidence (21) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Introducing GPT-6 Sol and Luna
singularity
Retrieved article excerpt

Open article · Retrieved 2026-09-22T18:24:17.863121+00:00

# Introducing GPT‑6 Sol and Luna

More ways to bring frontier intelligence into the work you do every day.

Loading…

Share

Earlier this month, we introduced [GPT‑6 Astra](https://openai.com/index/gpt-6-astra/), the most intelligent and aligned model in the world. While the most demanding and important projects still call for Astra’s full depth, work happens at different scales, rhythms, and budgets.

That’s why we’re expanding the GPT‑6 universe with **GPT‑6 Sol** and **GPT‑6 Luna.** GPT‑6 Astra introduced a new generation of intelligence—these models help distribute the benefits of that intelligence by advancing the frontier on cost efficiency. We trained GPT‑6 Sol and Luna with similar methods as GPT‑6 Astra, bringing the advances behind Astra’s state-of-the-art performance in professional work, factuality, coding, computer use, and alignment to faster, more affordable models.

The GPT‑6 models lead across the cost–intelligence curve, combining exceptional capabilities at every tier with infrastructure that delivers them efficiently at scale. Improvements in caching and inference let us serve these models at lower cost, and we’re passing those savings directly on to users and customers by **reducing API prices for Sol and Luna by 50%** compared with their GPT‑5.6 promotional pricing. Together, these improvements make advanced AI practical for more everyday tasks and applications at scale.

### GPT‑6 API pricing

|  |  |  |  |
| --- | --- | --- | --- |
| **Model** | **Input** | **Output** | **Price reduction** |
| **GPT‑6 Sol** vs. GPT‑5.6 Sol | $4 → **$2** | $20 → **$10** | 50% cheaper |
| **GPT‑6 Luna** vs. GPT‑5.6 Luna | $0.20 → **$0.10** | $1.20 → **$0.50** | 50% cheaper |

*Prices are per 1 million tokens.*

**GPT‑6 Astra** continues to be our best model across the board. Choose it when you want the best results and an uncompromising experience.

## A step up across the model family

GPT‑6 Sol and Luna bring intelligence upgrades and cost efficiency to the models you already know and use across capabilities most useful for getting complex work done.

### Professional work

GPT‑6 Sol can take on difficult work tasks while giving you more room to iterate with higher usage limits and lower cost, offering more intelligence and better results versus similarly priced competitor models.

On **AutomationBench,** a test of business workflows across apps, GPT‑6 Sol at xhigh effort outperforms Claude Opus 5 at max effort at just 9% of Opus 5’s cost per task. At high effort, GPT‑6 Luna improves on its predecessor by 5.4 percentage points at 58% lower cost per task.

*In* [*AutomationBench 1.0.6*⁠(opens in a new window)](https://zapier.com/benchmarks)*, AI agents are tested on end-to-end workflows using 47 tools across sales, marketing, operations, support, finance, and HR. The datapoint for Claude Fable 5.1 understates its actual cost, as it omits the cost of the Opus 5 fallbacks, which occurred on ~40% of tasks.*

GPT‑6 Sol also exceeds Claude Fable 5.1 at far lower cost, and even bests low-effort GPT‑6 Astra.

| **Model (and effort)** | **Score** | **Cost per task** |
| --- | --- | --- |
| GPT‑6 Sol (xhigh) | 33.2% | $0.27 |
| GPT‑6 Astra (low) | 30.3% | **3.9x** GPT‑6 Sol |
| Claude Opus 5 (max) | 26.9% | **11.1x** GPT‑6 Sol |
| Claude Fable 5.1 w/ Opus 5 Fallback (max) | 31.4% | **>8.9x** GPT‑6 Sol  *(fallback cost not reported)* |

On **Agents’ Last Exam**, which evaluates agents on complex professional workflows, GPT‑6 Sol at max effort scores 56.4%, above Claude Opus 5’s highest score in the evaluation at 60% lower cost per task.

*In* [*Agents’ Last Exam V1*⁠(opens in a new window)](https://agents-last-exam.org/)*, AI agents are evaluated on long-horizon, economically valuable tasks spanning 55 sub-industries, covering most major fields of professional work performed on a computer.*

### Factuality

The usefulness of an answer depends on getting the facts right, and we’re continuing to make progress on factual reliability. On our internal factuality evaluation, which is based on de-identified real-world conversations where users flagged mistakes by our models, GPT‑6 Sol makes about half as many mistakes as its predecessor, approaching Astra-level reliability at much lower cost. GPT‑6 Luna also improves substantially; at higher effort levels it matches GPT‑5.6 Sol at about a hundredth its cost.

*Here we evaluate factuality on de-identified ChatGPT conversations where users had flagged a factual error from a prior model. These error-inducing conversations are not representative of typical usage, where factual errors are more rare. Scores are not controlled for length; however, our verbosity sweeps showed almost no dependence on answer length.*

### Coding

This year, coding agents have begun tackling tasks with more complexity, scope, and duration than ever before. At OpenAI, our internal usage has grown exponentially. Valued at API prices, daily token usage has exceeded $600 for the median researcher and $7,000 for researchers at the 90th percentile ([Research acceleration: The view inside OpenAI⁠](https://openai.com/index/research-acceleration-view-inside-openai/)). As coding agents take on longer and more demanding tasks, the cost of sustained use matters more. GPT‑6 Sol and Luna combine strong coding performance with lower API prices, giving developers more room to iterate and teams the confidence to be more ambitious about what they ask Codex to take on.

On **FrontierCode**, which evaluates whether coding agents produce changes ready to merge into real codebases, GPT‑6 Sol improves substantially over GPT‑5.6 Sol, and is able to match Claude Fable 5.1 xhigh at much lower cost.

*In* [*FrontierCode 1.1 Main*⁠(opens in a new window)](https://cognition.com/frontiercode)*, AI agents write code that’s graded not only on correctness but also “mergeability”: e.g., test quality, scope discipline, code style, and adherence to codebase standards.*

On **DeepSWE v1.1,** which tests performance on complex software-engineering tasks in real codebases, GPT‑6 Sol at max effort scores 68.8%, within 1.1 percentage points of Claude Fable 5’s highest score in the evaluation—69.9% at xhigh effort—at approximately 80% lower cost per task.

GPT‑6 Luna at max effort scores 66.6%, comparable to Claude Opus 5 and Fable 5 at medium effort. In these comparisons, Luna costs 93% less per task than Opus 5 and 96% less than Fable 5.

*In* [*DeepSWE 1.1*⁠(opens in a new window)](https://deepswe.datacurve.ai/)*, AI agents solve original, long-horizon software engineering tasks.*

### Computer use

While GPT‑6 Astra remains the world’s best model for computer use, GPT‑6 Sol and Luna offer more cost-efficient performance than their predecessors. On **OSWorld 2.0 offline**, GPT‑6 Sol at xhigh effort achieves a similar score to Claude Opus 5 at medium effort—60.5% versus 60.3%—at approximately 80% lower cost per task. GPT‑6 Luna (max) is able to exceed GPT‑5.6 Sol (medium) at one tenth of its cost.

*In* [*OSWorld 2.0*⁠(opens in a new window)](https://osworld-v2.xlang.ai/)*, AI agents attempt long-horizon computer-use workflows spanning everyday and professional tasks. We report the partial reward on the offline set from the v2026.08.08 release.*

### Collaboration style

We’ve also brought GPT‑6 Astra’s improved communication style to Sol and Luna, which we think will be especially noticeable in technical and coding conversations. Expect to see more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall without losing substance.

Prompt

Website’s looking clean! Could we use a Bento Box design style and add a slider between the pages in the top right? You may need to dive into the React..

GPT-5.6 Sol

Absolutely. I can get the bento feel and the top-right sliding page switcher without pulling in React—the interaction is small enough to stay fast and smooth in the current site. I’m reshaping all four pages around that system now.

---

Done — open the live site.

It now has a colorful bento-box layout across all four pages and an animated top-right Home / Play / Place / Taste slider. React wasn’t necessary; native page transitions keep it lightweight and smooth.

The built-in image tool created the food artwork. Prompt: “An energetic gouache-and-ink editorial illustration celebrating dumplings and spicy noodles, with bright colors, steam, spice, and no text or logos.”

GPT-6 Sol

I’ll reshape the pages into a bento grid and turn the top-right links into a sliding page switcher. I’ll keep the existing site and check whether that interaction needs React before changing its setup.

---

The updated site is live. Every page now has a playful bento layout, and the top-right page switcher slides between sections. I checked it on desktop and narrow mobile screens, including browser back navigation. The existing site didn’t need React for this.

*Although style is subjective, we prefer GPT‑6 Sol’s reply here. It doesn’t jump to conclusions as quickly, spends less time reiterating details that might be obvious to the asker (e.g., that the website has four pages), uses less vague language (e.g., “bento feel”, “reshaping… around that system”), is more forthcoming with what it did and didn’t check, and doesn’t unnecessarily share implementation details like its image tool prompt.*

## Improving caching for agents and long conversations

Alongside lower token prices, we’re helping developers building on GPT‑6 save more on the context their applications reuse. We’ve improved prompt caching for GPT‑6 to deliver higher cache hit rates by default, helping agents reuse more context, respond faster, and benefit from discounts of 90% on cached input-token reads.

Developers also have more ways to measure and optimize their caching performance:

- **Monitor and diagnose.** The [Prompt Caching Dashboard⁠(opens in a new window)](https://platform.openai.com/usage?usage_section=prompt-caching) shows how much input is cached and how that changes over time. The [diagnostics tool⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics) helps explain missed opportunities for caching and what to fix.
- **Adjust reasoning effort and tool availability without breaking cache.** Increase [reasoning effort⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/reasoning#change-reasoning-mid-conversation) for harder tasks or lower it for simpler follow-ups, and [enable or disable tools⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching#how-to-optimize-prompt-caching) as your agent’s needs change. Both controls now preserve earlier context for cache reuse.
- **Optimize which prefixes get cached.** Explicit breakpoints let developers choose where cached prompt prefixes end. This gives developers more control over cache reuse and can improve performance.

GitHub reports that, over the past several months, these improvements have reduced the share of prompt tokens requiring fresh processing by more than 50% across billions of requests to OpenAI models, helping Copilot respond faster.

## Continuing to improve alignment

GPT‑6 Sol and Luna build on the alignment work introduced with Astra, our most aligned model to date. In our alignment evaluations, both Sol and Luna show improvements over their GPT‑5.6 counterparts, including lower rates of misleading claims about their coding work.

The evaluations below deliberately test challenging situations and do not measure failure rates in typical use. See the [system card⁠(opens in a new window)](https://deploymentsafety.openai.com/gpt-6-astra) for the full results.

## Availability

GPT‑6 Sol and GPT‑6 Luna are available in ChatGPT Work and Codex starting today for all Plus, Pro, Business, Enterprise, and Edu users. Free and Go users can access GPT‑6 Luna in the desktop app. These models are not yet available in Chat. In the OpenAI API, they are 
DemiPixel1331313
🟠 redditGPT 6 Sol/Luna available on Codex/Work
singularity
Retrieved article excerpt

Open article · Retrieved 2026-09-22T18:24:16.402189+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
ees-h2710
🟠 redditOpenAI launches GPT-6 Sol + Luna: Astra intelligence moves down the cost curve
OpenAI
Retrieved article excerpt

Open article · Retrieved 2026-09-22T19:23:33.605954+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
etherd0t10116
🟠 redditGPT-6 Sol surpasses Claude Opus 5 on Agents’ Last Exam at 60% lower cost
singularity
Retrieved article excerpt

Open article · Retrieved 2026-09-22T22:25:24.932975+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
141_13379331
🟠 redditGPT 6 Sol and Luna on benchmarks
singularity
krizzalicious494510
🟠 redditGPT-6 Sol scores 33.2% on AutomationBench, beating Opus 5 while costing ~11× less per task
singularity
141_1337317
🟠 redditGPT 6 Sol and Luna at half the GPT 5.6 prices
singularity
No_Hovercraft623913215
🟠 redditGPT 6 Sol and Luna Prices
singularity
Ok_Barracuda_1161173
🟧 hnGPT-6 Sol and LunaOfficialTurkey1723819
🟧 hnGPT-6 Sol (Max) Intelligence, Performance and Price Analysis
Retrieved article excerpt

Open article · Retrieved 2026-09-22T21:23:08.354203+00:00

[Artificial Analysis](https://artificialanalysis.ai/)

K

OpenAI logo

[OpenAI](https://openai.com/)

•

[GPT-6 Sol](https://artificialanalysis.ai/models/releases/gpt-6-sol)max

•

Proprietary model

•

Released September 2026

# GPT-6 Sol (max) Intelligence, Performance & Price Analysis

Compare[Try it out](https://artificialanalysis.ai/microevals) [API Provider Benchmarks](https://artificialanalysis.ai/models/gpt-6-sol/providers)

### Model summary

#### [Intelligence](https://artificialanalysis.ai/models/gpt-6-sol#intelligence)Updated

#18 / 212

48

Artificial Analysis Intelligence Index

4 out of 4 units for Intelligence.

#### [Speed](https://artificialanalysis.ai/models/gpt-6-sol#speed)

#49 / 212

104.4

Output tokens per second

3 out of 4 units for Speed.

#### [Cost](https://artificialanalysis.ai/models/gpt-6-sol#price-cost)

In $2.00Out $10.00Cache Discount 90%

N/A

Cost per Intelligence Index task

Unknown out of 4 units for Cost.

#### [Verbosity](https://artificialanalysis.ai/models/gpt-6-sol#token-use)

#43 / 212

77M

Output tokens from Intelligence Index

2 out of 4 units for Verbosity.

### Comparison Summary

GPT-6 Sol (max) is amongst the leading models in intelligence and reasonably priced when comparing to other models of similar price. It's also faster than average and fairly concise. The model supports text and image input, outputs text, and has a 872k tokens context window.

GPT-6 Sol (max) scores 48 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median: 25). When evaluating the Intelligence Index, it generated 77M tokens, which is fairly concise in comparison to the median of 88M.

Pricing for GPT-6 Sol (max) is $2.00 per 1M input tokens (moderately priced, median: $2.00) and $10.00 per 1M output tokens (moderately priced, median: $10.00).

At 104 tokens per second, GPT-6 Sol (max) is faster than average (75).

### Technical specifications

|  |  |
| --- | --- |
| Reasoning | Yes This page shows the reasoning version of this model.  A non-reasoning variant may also exist. |
| Input modality | Supports: text and image |
| Output modality | Supports: text |
| Context window | 872k ~1308 A4 pages of size 12 Arial font |

### 212 models in this class

Metrics are compared against models of the same class:

- Non-reasoning models → compared only with other non-reasoning models
- Reasoning models → compared across both reasoning and non-reasoning
- Open weights models → compared only with other open weights models of the same size class:

- Tiny: ≤4B parameters
- Small: 4B–40B parameters
- Medium: 40B–150B parameters
- Large: >150B parameters

- Proprietary models → compared across proprietary and open weights models of the same price range, using a blended 3:1 input/output price ratio:

- <$0.15 per 1M tokens
- $0.15–$1 per 1M tokens
- >$1 per 1M tokens

Highlights

Updated

### [Intelligence](https://artificialanalysis.ai/models/gpt-6-sol#intelligence)

Artificial Analysis Intelligence Index · Higher is better

### [Speed](https://artificialanalysis.ai/models/gpt-6-sol#speed)

Output tokens per second · Higher is better

### [Cost per Task](https://artificialanalysis.ai/models/gpt-6-sol#price-cost)

Weighted average cost (USD) per Intelligence Index task · Lower is better

Prompt Options

## IntelligenceUpdated

### [Artificial Analysis Intelligence Index](https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index)

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

28 of 673 models

Add model from specific provider

### Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See [Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) for further details, including a breakdown of each evaluation and how we run them.

Open Weights / ProprietaryReasoning / Non-ReasoningText Only / Multimodal Inputs

### Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

28 of 673 models

Add model from specific provider

ProprietaryOpen WeightsOpen Weights (Commercial Use Restricted)

### Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See [Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) for further details, including a breakdown of each evaluation and how we run them.

### Open Weights

Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if commercial use is limited by conditions, and as 'Non-commercial' if the license prohibits commercial use.

## [Capability Indexes](https://artificialanalysis.ai/models/capabilities)

Measures the performance of models on specific capabilities and industries

Finance & AccountingStrategy & OpsLegalEngineeringEconomics

### [Artificial Analysis Finance & Accounting Index](https://artificialanalysis.ai/models/capabilities/finance-and-accounting)

Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better

28 of 170 models

Add model from specific provider

## [Benchmarks](https://artificialanalysis.ai/evaluations)

### Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better

CodingAgenticTool UsePrivate DatasetUser InteractionFinanceMedicalLegalIntelligence IndexLong ContextMultimodalInstruction FollowingFaithfulnessWritingBusiness[See more](https://artificialanalysis.ai/evaluations)

18 of 26 evaluations

28 of 673 models

Add model from specific provider

[AA-Briefcase v1.1](https://artificialanalysis.ai/evaluations/aa-briefcase)Updated

Agentic knowledge work, (Elo-500)/2000

[GDPval-AA v2.1](https://artificialanalysis.ai/evaluations/gdpval-aa)Updated

Agentic real-world work tasks, (Elo-500)/2000

[AutomationBench-AA](https://artificialanalysis.ai/evaluations/automationbench-aa)Updated

Agentic SaaS workflows

[Terminal-Bench 4.0](https://artificialanalysis.ai/evaluations/terminalbench-4-0)New

Agentic coding & terminal use

[SciCode](https://artificialanalysis.ai/evaluations/scicode)

Coding

[Humanity's Last Exam](https://artificialanalysis.ai/evaluations/humanitys-last-exam)

Reasoning & knowledge

[GDP.pdf](https://artificialanalysis.ai/evaluations/gdp-pdf)New

Professional document reasoning, All-pass

[CritPt](https://artificialanalysis.ai/evaluations/critpt)

Physics reasoning

[AA-Omniscience Accuracy](https://artificialanalysis.ai/evaluations/omniscience)

Knowledge

[AA-Omniscience Non-Hallucination Rate](https://artificialanalysis.ai/evaluations/omniscience)

1 - hallucination rate

[AA-LCR v1.1](https://artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning)

Long context reasoning

[Harvey LAB-AA](https://artificialanalysis.ai/evaluations/harvey-lab-aa)

Legal agentic work, criterion pass rate

[EnterpriseOps-Gym-AA](https://artificialanalysis.ai/evaluations/enterprise-ops-gym-aa)

Agentic business operations

[AA-AnalystAgent](https://artificialanalysis.ai/evaluations/aa-analyst-agent)

Quantitative analysis on spreadsheets & documents

[𝜏³-Banking](https://artificialanalysis.ai/evaluations/tau3-banking)

Agentic tool use

[ITBench-AA](https://artificialanalysis.ai/evaluations/itbench-aa)

Kubernetes incident root-cause analysis

[MMMU-Pro](https://artificialanalysis.ai/evaluations/mmmu-pro)

Visual reasoning

[MLCR-AA](https://artificialanalysis.ai/evaluations/mlcr-aa)New

Medical long context reasoning

### Intelligence Evaluation Relevance

While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.

### Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See [Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) for further details, including a breakdown of each evaluation and how we run them.

### AA-Briefcase v1.1Updated

AA-Briefcase EloAA-Briefcase Rubric Score (%)Analytical Quality & Presentation Elo

### AA-Briefcase Elo

AA-Briefcase v1.1 is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better

28 of 191 models

Add model from specific provider

### AA-Briefcase Elo

AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.

### AA-Omniscience

AA-Omniscience IndexAA-Omniscience AccuracyAA-Omniscience Hallucination Rate

### AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

28 of 548 models

Add model from specific provider

### AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

## Intelligence Index Comparisons

Intelligence Index vs. Cost per TaskIntelligence Index vs. Time per TaskIntelligence Index vs. Output SpeedIntelligence Index vs. End-to-End Response Time

### Intelligence Index vs. Cost per Intelligence Index Task

Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task

28 of 673 models

Most attractive quadrant

Pareto line

XiaomiOpenAIAnthropicSpaceXAIGoogleMetaZ AIDeepSeek

### Cost per Intelligence Index Task

Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.

### Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See [Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) for further details, including a breakdown of each evaluation and how we run them.

## Token Use

Output Tokens per TaskIntelligence Index vs. Output Tokens per TaskIntelligence Index Token UseIntelligence Index vs. Token Use

### Output Tokens per Intelligence Index Task

Weighted average number of output tokens used to run one task in the Artificial Analysis Intelligence Index

28 of 673 models

AnswerReasoning
theanonymousone110
🟧 echo.blogAnnounces Sol and Luna availability in Work, Codex, and the API, lower token pricing, improved prompt caching, and claimed benchmark cost-peOpenAI——
🟠 reddit6 Sol Nerfed?
singularity
Ok_Flamingo_301231
🟠 redditGPT‑6 Sol and Luna: Cheaper, but Worse Where It Matters
OpenAI
AirportEither2456820
🟠 redditGPT-6 Sol is NOT the replacement of GPT-5.6 Sol...
OpenAI
Jarr11477151
🟠 redditSol-6 Be super careful
OpenAI
Babayaga166423461
🟠 redditSol and Luna feel weaker after this week's update, but Astra's usage limits seem massively improved
OpenAI
Shay_Solomon196
🟠 redditOpenAI launches GPT-6.1 Sol
OpenAI
ethotopia685149
🟠 redditIntroducing 6.1 SOL
OpenAI
BitterAd641911735
🟠 redditOpus 5.5 still leads on AA intelligence, but GPT-6.1 Sol shifts the cost frontier
OpenAI
rhiever141
🟠 redditOpus 5.5 still leads on AA intelligence, but GPT-6.1 Sol shifts the cost frontier
ClaudeAI
rhiever21960
🟠 redditI don't understand the constant bashing of GPT 6/6.1 Sol's performance relative to 5.6 and Astra. Why are we ignoring the elephant in the room, the absurd price + efficiency gains? It costs half as much as Sonnet 5.5, 1/3rd of 5.6 Sol, 1/5th of Opus 5.5, and 1.7th of Astra.
singularity
that_90s_guy153102

Interpretation history

Decision trace