2026-10-11 16:37 UTC

Google claims its newly announced Gemini 4 Argon delivers frontier-leading real-work capability — SOTA DeepSWE v1.1 (77.9%), Vals Index and CWE-bench leads, 1M-token input and output, $2/$10 per-million pricing, phased rollout to trusted cyber defenders under US-government pre-release evaluation — and whether that holds in hands-on coding and agent use (early counter-signals: Artificial Analysis #8/223 intelligence, Bloomberg-reported internal doubts on real coding work) decides whether it displaces GPT-6 Astra and Opus 5.5 as a default for agent workloads.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highfrontier-model-releases gemini-4-argon model-economics agent-evaluation ai-governanceGoogle DeepMindKoray KavukcuogluTulsee DoshiSundar PichaiWiz
Surfaced 2026-09-30T23:54:08Z — Announces Gemini 4 Argon, 'our next era of frontier intelligence': frontier performance across real-world software engineering, enterprise k — Google claims its newly announced Gemini 4 Argon delivers frontier-leading real-work capability — SOTA DeepSWE v1.1 (77.9%), Vals Index and CWE-bench leads, 1M-token input and output, $2/$10 per-million pricing, phased rollout to trusted cyber defenders under US-government pre-release evaluation — and whether that holds in hands-on coding and agent use (early counter-signals: Artificial Analysis #8/223 intelligence, Bloomberg-reported internal doubts on real coding work) decides whether it displaces GPT-6 Astra and Opus 5.5 as a default for agent workloads.

What is this?

Google DeepMind announced Gemini 4 Argon on 2026-09-30 as a frontier flagship claiming state-of-the-art real-work results (77.9% DeepSWE v1.1, leads on the Vals Index and CWE-bench), 1M-token input and output, and $2/$10 per-million 'introductory' pricing — with phased access that gates cyber capability to vetted defenders via 'Fairwind' under US-government pre-release evaluation rather than opening general availability. The supplied snippets do not directly document the announcement itself: the closest hit (Dataconomy, Sept 21) reports the pre-release leak — an Argon Arena listing near 88% DeepSWE at suspected $2.25/$11.25 pricing — consistent with the case's recorded pre-ship figures (88.7%/2M context) and its leak-vs-ship downgrade to 77.9%/1M. The counter-picture that Argon is benchmaxxed — Artificial Analysis at #8/223, arena-style reviews ranking it below Fable 5/Opus 5.5/GPT-6 Sol, and Bloomberg-reported internal doubts it struggles with real coding work — rests on the case's own evidence titles, not on the supplied web hits. The snippets do corroborate the surrounding market: Google's Sept 2 Gemini 3.8 Flash launch already set the intro-price-then-raise pattern ($0.75/$3.75 doubling on Jan 1, 2027) and a gated 'Flash Cyber' variant, rival specs (GPT-6 Astra 74.1% DeepSWE, Opus 5.5 leading agentic coding) frame exactly the displacement question the case poses, and one comparison site independently voices the case's core lens — '13x more per output token for a lead of twelve points.'

Why it matters to Scott

Converges on two of Scott's core positions at once, with consequential parties: Bloomberg-sourced Google-employee admissions that Argon 'crushes benchmarks but struggles with real work' are a dated, vendor-internal receipt for his benchmarking-the-wrong-unit thesis, and AA's Coding Agent Index — Argon's token-price advantage inverting to a ~5.6x per-task cost disadvantage vs Sol 6.1 — is an independent instrument measuring exactly the token-price-vs-cost-per-task unit his Agent Token Manifesto argues. It also bears on live decisions: his LiteLLM cost-tier routing keeps Argon out of the coding tier, and the untested 1M-output claim is precisely the residence substrate his Inference Field ebook would interrogate first if the Fairwind gate opens.
ip:concept.benchmarking-the-wrong-unitip:concept.model-plus-harness-benchmark-unitip:concept.ai-unit-economicsdev:project.llmreportdev:concept.cost-tiered-llm-routingip:source.the-inference-field-ebookradar:concept.agent-evaluationradar:concept.benchmark-integrityradar:concept.agent-harnessesradar:concept.inference-economicsradar:concept.model-pricingradar:sonnet-55-release-economicsradar:claude-sonnet-55-token-economicsradar:openai-gpt6-sol-luna-releaseradar:white-house-gold-eagle-frontier-access
queries asked of Scott's wikis
  • benchmark vs real-work gap benchmaxxing eval realism position
  • token price vs cost-per-task agent economics
  • harness choice changes agent benchmark scores CLI
  • gated model access cyber restrictions government pre-release evaluation
  • 1M output context vs RAG for agent workloads
  • introductory pricing frontier model price war pattern

Measured heat

now 0 pts/hpeak 1576 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 290h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-29 14:00⭐ origin echo-reconstructedAnnounces Gemini 4 Argon, 'our next era of frontier intelligence': frontier performance across real-world software engineering, enterprise k
Google (Koray Kavukcuoglu, SVP Google DeepMind and Chief AI Architect) on blog (echo) · attributed from hn.story.49913571, reddit.post.1wufk1u, reddit.post.1wufl5b, reddit.post.1wuga3s, hn.story.49914236, reddit.post.1wufsth, reddit.post.1wufw1f, reddit.post.1wufikm, reddit.post.1wug9tb, reddit.post.1wug4vf, reddit.post.1wug38q, reddit.post.1wufufo, reddit.post.1wufe9n, reddit.post.1wufdet
—
09-30 20:04first on hacker news · published · +30.1hGemini 4 Argon
bradleyg223
—
09-30 20:08first on r/singularity · published · +30.1hGoogle’s unreleased Gemini 4 Argon may have just leaked—and it tops 12 of 18 benchmarks against Fable 5.1, Opus 5.5 and GPT-6 Astra, including 19.6% vs GPT-6 Astra’s 5.4% on autonomous legal work
141_1337
—
09-30 20:13first on r/artificial · published · +30.2hGoogle cooked OpenAI and Anthropic with Gemini 4 Argon
DataRemarkable7093
—
09-30 20:15first on r/ClaudeAI · published · +30.3hGemini 4 is out: The competition has woken up
monsieurcliffe
—
09-30 20:16first on r/OpenAI · published · +30.3hWake up babe, it's Gemini 4 fr
MrTimeHacker1
—
09-30 20:04amplified on hacker news 👑hn.story.49913571
bradleyg223
peak 1704 · 1188 comments · 41% of case engagement
09-30 20:08amplified on r/singularityreddit.post.1wufdet
141_1337
peak 228 · 96 comments · 3% of case engagement
09-30 20:09amplified on r/singularityreddit.post.1wufe9n
Every_Foundation5197
peak 131 · 20 comments · 1% of case engagement
09-30 20:13amplified on r/artificialreddit.post.1wufikm
DataRemarkable7093
peak 155 · 85 comments · 2% of case engagement
09-30 20:15amplified on r/ClaudeAIreddit.post.1wufk1u
monsieurcliffe
peak 1199 · 278 comments · 12% of case engagement
09-30 20:16amplified on r/OpenAIreddit.post.1wufl5b
MrTimeHacker1
peak 479 · 87 comments · 4% of case engagement
20 more amplifiers in ainews.case_chain
09-30 21:20our radar first saw it · +31.3hdiscovery anchor: hn.story.49913571—
09-30 21:42reached heat=high · +31.7h · via ledger——
pace: p99 vs 1188 stories at the 168h mark (now 290h old) — ahead of openai-millennium-maths-claim (1.3x), behind fable-51-multimodal-coding-cost (0.9x)

Evidence (27) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnGemini 4 Argon
Retrieved article excerpt

Open article · Retrieved 2026-09-30T21:38:10.861909+00:00

# Gemini 4 Argon: our next era of frontier intelligence

Sep 30, 2026

|

- [x.com](https://twitter.com/intent/tweet?text=Gemini%204%20Argon%3A%20our%20next%20era%20of%20frontier%20intelligence%20%40google&url=https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/)
- [Facebook](https://www.facebook.com/sharer/sharer.php?caption=Gemini%204%20Argon%3A%20our%20next%20era%20of%20frontier%20intelligence&u=https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/)
- [LinkedIn](https://www.linkedin.com/shareArticle?mini=true&url=https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/&title=Gemini%204%20Argon%3A%20our%20next%20era%20of%20frontier%20intelligence)
- Mail
- Copy link

Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.

---

[Koray Kavukcuoglu

SVP, Google DeepMind and Chief AI Architect, Google](https://blog.google/authors/koray-kavukcuoglu/)

Share

- [x.com](https://twitter.com/intent/tweet?text=Gemini%204%20Argon%3A%20our%20next%20era%20of%20frontier%20intelligence%20%40google&url=https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/)
- [Facebook](https://www.facebook.com/sharer/sharer.php?caption=Gemini%204%20Argon%3A%20our%20next%20era%20of%20frontier%20intelligence&u=https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/)
- [LinkedIn](https://www.linkedin.com/shareArticle?mini=true&url=https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/&title=Gemini%204%20Argon%3A%20our%20next%20era%20of%20frontier%20intelligence)
- Mail
- Copy link

---

Stylized promotional blog key art graphic with modern editorial branding and the text "Gemini 4 Argon"

Your browser does not support the audio element.

Listen to article

[[duration]] minutes

This content is generated by Google AI. Generative AI is experimental



Voice






Speed

Voice

Speed
0.75X
1X
1.5X
2X

Read AI-generated summary

- Google’s new Gemini 4 Argon model brings advanced reasoning to complex, long-horizon professional tasks.
- The model features an industry-leading 1 million token limit for deep, multi-step problem solving.
- It excels at coding, financial research, legal drafting, and autonomous cybersecurity vulnerability patching.
- Argon is currently rolling out to trusted cyber defenders through the Fairwind Program.
- Google is prioritizing safety and rigorous testing before a wider release to the public.

Summaries were generated by Google AI. Generative AI is experimental.

In this article



---

Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our [Fairwind Program](https://deepmind.google/fairwind-program/). Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Google. It delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.

Safely releasing frontier capabilities at this level requires a phased approach. We are actively engaged in the U.S. government’s voluntary process for pre-release model access while we gradually expand access. We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.

Argon will launch at an introductory price
[1](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/#footnote-1)
of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.

## Changing how we work and build at Google

Gemini 4 Argon is already powering our internal workflows, with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality. It’s helping teams build faster and push the boundaries of engineering productivity and accelerating breakthroughs:

- **Quantum algorithmic optimization:** Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.

- **Memory efficiency:** A team of Argon agents analyzed fleet-wide profiling telemetry to autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.

- **Large Scale Codebase Migrations and Optimizations:** Argon agents are working on migrating C/C++ codebases to Rust across Google—scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel. Given the criticality of many of these systems, such large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production.  
    
  For libgav1, Google's open source software for decoding video, Argon agents took an existing Rust port and replaced 32K lines of SIMD code by running many rounds of profile-guided experiments, studying the compiler's output, producing safe Rust so the compiler would vectorize it automatically. The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.

## Working harder on your most complex problems

To support Gemini 4 Argon’s capabilities across longer, more complex use cases, we are significantly expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens. When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go.



a benchmark chart showing Gemini 4 Argon capabilities

## Enabling coding and enterprise workflows across domains

Gemini 4 Argon’s capabilities across coding, reasoning, and multimodality and its ability to sustain long, multi-step tasks enable it to excel across a range of enterprise workflows.

Google engineers have been using Argon for their daily tasks, from everyday debugging to large-scale codebase migrations and algorithm designs. It sets a new state of the art on DeepSWE v1.1 (77.9%), which measures a model’s performance in real-world long-horizon software engineering tasks.

Beyond coding, Argon is the leading model on the [Vals Index](https://www.vals.ai/benchmarks/vals_index), which measures economic impact across finance, coding, legal, and tax work, with every sector weighted by its contribution to U.S. GDP. We see similarly leading performance across other domain specific evaluations, like Vals Finance Agent v2 (multi-step financial research) and Harvey’s Legal Agent Benchmark (legal research and drafting). On AutomationBench, Zapier’s benchmark measuring end-to-end execution across core business functions, Argon ranks #1 with a score of 51.3%.

Argon is also uniquely strong when knowledge work requires visual understanding. It’s able to drive professional chart analysis, identify details from long videos, and take action based on a series of documents. For example, on LVBench, which measures long video understanding, Argon is state of the art with a score of 91.7%.

DeepSwe evaluation chart



Vals index chart



Val's finance benchmark chart



Harveys benchmark chart



automation bench chart

## Leading in defensive cybersecurity

To better equip cyber defenders for the new era of cyberattacks, we trained Gemini 4 Argon to be highly capable at cybersecurity defense. Argon can autonomously find, validate, and patch critical software vulnerabilities. For trusted defenders and our own internal teams at Google, we’ll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities.

[Wiz](https://www.wiz.io/) is already using Argon for cybersecurity defense through its [Scan for Good](https://www.wiz.io/scan-for-good) initiative – a program dedicated to protecting critical public infrastructure for free by finding and remediating high-risk exposures. In an early demonstration of its impact, the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, identifying a severe risk that previous frontier models had missed.

On [CWE-bench v1](https://cwe-bench.com/), which evaluates the model’s ability to remediate security vulnerabilities, Argon ties for first place with a top score of 68%, building on 3.8 Flash Cyber’s frontier performance on [CWE-bench v0](https://cwe-bench.com/?v=v0).



CWE benchmark chart

Gemini 4 Argon demonstrates impressive leaps in vulnerability discovery over 3.8 Flash Cyber. For example:

- On Google’s internal comprehensive vulnerability benchmark, Argon uncovered a wide range of exposures across complex codebases spanning 20 programming languages.
- On Wiz’s internal black-box penetration testing benchmark, which tests a model’s ability to analyze live web systems without source code, Argon outperforms 3.8 Flash Cyber in discovering the attack surface, identifying vulnerabilities, and producing proof-of-concept evidence to validate them.



security vulnerabilities chart

## Strengthening frontier safeguards before broad availability

Before rolling out Gemini 4 Argon broadly, we’re continuing to strengthen critical frontier safeguards across four main areas:

**Defending against misuse:** To prevent bad actors from using Argon for cyber or chemical, biological, radiological, and nuclear (CBRN) attacks, the model is designed to refuse harmful requests while preserving legitimate, dual-use scientific research, as per our [Frontier Safety Framework](https://deepmind.google/blog/strengthening-our-frontier-safety-framework/). We are strengthening the robustness of our safeguards for this launch, including improving our techniques to monitor the model’s [internal activations](https://arxiv.org/abs/2601.11516) to spot misuse. These safeguards underwent robustness testing by internal and external red teams using a combination of manual and automated attack methods.

**Defending against prompt injection attacks:** Argon is also our most resilient model yet against indirect prompt injections, where malicious instructions or context are used to hijack a model’s behavior. These are complex attacks that require constant vigilance and multiple layers of defense. Through automated red teaming and adversarial training, Gemini 4 Argon is leading in prompt injection robustness on the Gray Swan’s Indirect Prompt Injection (IPI) benchmark.



Gray Swan evaluation

**Monitoring for misalignment:** In order to prevent Argon from stepping out of bounds to try to accomplish a task in a way that goes beyond the user’s intentions, we are deploying misalignment mitigations that monitor Argon’s chain-of-thought and actions and stop execution when necessary.

We used a similar system to monitor our training runs and send alerts to a dedicated incident response team, taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We [strongly encourage the rest of the industry](https://institute.deepmind.com/essays/the-case-for-reasoning-transparency/) to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.

**Hardeni
bradleyg22317041188
🟠 redditGemini 4 is out: The competition has woken up
ClaudeAI
Retrieved article excerpt

Open article · Retrieved 2026-09-30T21:38:08.844028+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
monsieurcliffe1198278
🟠 redditWake up babe, it's Gemini 4 fr
OpenAI
Retrieved article excerpt

Open article · Retrieved 2026-09-30T21:38:04.445609+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
MrTimeHacker147887
🟠 redditGemini 4 Crushes Benchmarks, But Google Employees State The Model Struggles With Real Work
singularity
Neurogence17390
🟧 hnGemini 4 Argon (High): Intelligence, Performance and Price Analysis
Retrieved article excerpt

Open article · Retrieved 2026-09-30T21:38:12.095195+00:00

[Artificial Analysis](https://artificialanalysis.ai/)

K

Google logo

[Google](https://artificialanalysis.ai/models/creators/google)

•

Proprietary model

•

Released September 2026

# Gemini 4 Argon (High) Intelligence, Performance & Price Analysis

Compare[Try it out](https://artificialanalysis.ai/microevals) [API Provider Benchmarks](https://artificialanalysis.ai/models/gemini-4-argon/providers)

### Model summary

#### [Intelligence](https://artificialanalysis.ai/models/gemini-4-argon#intelligence)Updated

#8 / 223

53

Artificial Analysis Intelligence Index

4 out of 4 units for Intelligence.

#### Speed

N/A

Output tokens per second

Unknown out of 4 units for Speed.

#### [Cost](https://artificialanalysis.ai/models/gemini-4-argon#price-cost)

#77 / 223

In $2.00Out $10.00Cache Discount 95%

$1.99

Cost per Intelligence Index task

3 out of 4 units for Cost.

#### [Verbosity](https://artificialanalysis.ai/models/gemini-4-argon#token-use)

#74 / 223

110M

Output tokens from Intelligence Index

3 out of 4 units for Verbosity.

### Comparison Summary

Gemini 4 Argon (High) is amongst the leading models in intelligence and reasonably priced when comparing to other models of similar price. The model supports text and image input, outputs text, and has a 1M tokens context window.

Gemini 4 Argon (High) scores 53 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median: 26). When evaluating the Intelligence Index, it generated 110M tokens, which is somewhat verbose in comparison to the median of 82M.

Pricing for Gemini 4 Argon (High) is $2.00 per 1M input tokens (moderately priced, median: $2.00) and $10.00 per 1M output tokens (moderately priced, median: $10.00). On average, it costs $1.99 per task to evaluate Gemini 4 Argon (High) on the Intelligence Index.

### Technical specifications

|  |  |
| --- | --- |
| Reasoning | Yes This page shows the reasoning version of this model.  A non-reasoning variant may also exist. |
| Input modality | Supports: text and image |
| Output modality | Supports: text |
| Context window | 1M ~1500 A4 pages of size 12 Arial font |

### 223 models in this class

Metrics are compared against models of the same class:

- Non-reasoning models → compared only with other non-reasoning models
- Reasoning models → compared across both reasoning and non-reasoning
- Open weights models → compared only with other open weights models of the same size class:

- Tiny: ≤4B parameters
- Small: 4B–40B parameters
- Medium: 40B–150B parameters
- Large: >150B parameters

- Proprietary models → compared across proprietary and open weights models of the same price range, using a blended 3:1 input/output price ratio:

- <$0.15 per 1M tokens
- $0.15–$1 per 1M tokens
- >$1 per 1M tokens

Highlights

Updated

### [Intelligence](https://artificialanalysis.ai/models/gemini-4-argon#intelligence)

Artificial Analysis Intelligence Index · Higher is better

Not publicly available

### Speed

Output tokens per second · Higher is better

### [Cost per Task](https://artificialanalysis.ai/models/gemini-4-argon#price-cost)

Weighted average cost (USD) per Intelligence Index task · Lower is better

Not publicly available

Prompt Options

## IntelligenceUpdated

### [Artificial Analysis Intelligence Index](https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index)

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

25 of 687 models

Add model from specific provider

Not publicly available

### Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See [Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) for further details, including a breakdown of each evaluation and how we run them.

Open Weights / ProprietaryReasoning / Non-ReasoningText Only / Multimodal Inputs

### Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

25 of 687 models

Add model from specific provider

Not publicly available

ProprietaryOpen WeightsOpen Weights (Commercial Use Restricted)

### Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See [Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) for further details, including a breakdown of each evaluation and how we run them.

### Open Weights

Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if commercial use is limited by conditions, and as 'Non-commercial' if the license prohibits commercial use.

## [Capability Indexes](https://artificialanalysis.ai/models/capabilities)

Measures the performance of models on specific capabilities and industries

Finance & AccountingStrategy & OpsLegalEngineeringEconomics

### [Artificial Analysis Finance & Accounting Index](https://artificialanalysis.ai/models/capabilities/finance-and-accounting)

Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better

25 of 194 models

Add model from specific provider

Not publicly available

## [Benchmarks](https://artificialanalysis.ai/evaluations)

### Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better

CodingAgenticTool UsePrivate DatasetUser InteractionFinanceMedicalLegalIntelligence IndexLong ContextMultimodalInstruction FollowingFaithfulnessWritingBusiness[See more](https://artificialanalysis.ai/evaluations)

18 of 27 evaluations

25 of 687 models

Add model from specific provider

Not publicly available

[AA-Briefcase v1.1](https://artificialanalysis.ai/evaluations/aa-briefcase)Updated

Agentic knowledge work, (Elo-500)/2000

[GDPval-AA v2.1](https://artificialanalysis.ai/evaluations/gdpval-aa)Updated

Agentic real-world work tasks, (Elo-500)/2000

[AutomationBench-AA](https://artificialanalysis.ai/evaluations/automationbench-aa)Updated

Agentic SaaS workflows

[Terminal-Bench 4.0](https://artificialanalysis.ai/evaluations/terminalbench-4-0)New

Agentic coding & terminal use

[SciCode](https://artificialanalysis.ai/evaluations/scicode)Under review

Coding

[Humanity's Last Exam](https://artificialanalysis.ai/evaluations/humanitys-last-exam)

Reasoning & knowledge

[GDP.pdf](https://artificialanalysis.ai/evaluations/gdp-pdf)New

Professional document reasoning, All-pass

[CritPt](https://artificialanalysis.ai/evaluations/critpt)Under review

Physics reasoning

[AA-Omniscience Accuracy](https://artificialanalysis.ai/evaluations/omniscience)

Knowledge

[AA-Omniscience Non-Hallucination Rate](https://artificialanalysis.ai/evaluations/omniscience)

1 - hallucination rate

[AA-LCR v1.1](https://artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning)

Long context reasoning

[Harvey LAB-AA](https://artificialanalysis.ai/evaluations/harvey-lab-aa)

Legal agentic work, criterion pass rate

[EnterpriseOps-Gym-AA](https://artificialanalysis.ai/evaluations/enterprise-ops-gym-aa)

Agentic business operations

[Terminal-Bench-Science 0.1](https://artificialanalysis.ai/evaluations/terminal-bench-science)New

Agentic scientific research workflows in a terminal

[AA-AnalystAgent](https://artificialanalysis.ai/evaluations/aa-analyst-agent)

Quantitative analysis on spreadsheets & documents

[ITBench-AA](https://artificialanalysis.ai/evaluations/itbench-aa)

Kubernetes incident root-cause analysis

[MMMU-Pro](https://artificialanalysis.ai/evaluations/mmmu-pro)

Visual reasoning

[MLCR-AA](https://artificialanalysis.ai/evaluations/mlcr-aa)

Medical long context reasoning

### Intelligence Evaluation Relevance

While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.

### Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See [Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) for further details, including a breakdown of each evaluation and how we run them.

### AA-Briefcase v1.1Updated

AA-Briefcase EloAA-Briefcase Rubric Score (%)Analytical Quality & Presentation Elo

### AA-Briefcase Elo

AA-Briefcase v1.1 is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better

25 of 211 models

Add model from specific provider

Not publicly available

### AA-Briefcase Elo

AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.

### AA-Omniscience

AA-Omniscience IndexAA-Omniscience AccuracyAA-Omniscience Hallucination Rate

### AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

25 of 562 models

Add model from specific provider

Not publicly available

### AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

## Intelligence Index Comparisons

Intelligence Index vs. Cost per TaskIntelligence Index vs. Time per TaskIntelligence Index vs. Output SpeedIntelligence Index vs. End-to-End Response Time

### Intelligence Index vs. Cost per Intelligence Index Task

Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task

25 of 687 models

Most attractive quadrant

Pareto line

GoogleXiaomiOpenAIAnthropicSpaceXAIMetaZ AIDeepSeek

### Cost per Intelligence Index Task

Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.

### Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See [Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) for further details, including a breakdown of each evaluation and how we run them.

## Token Use

Output Tokens per TaskIntelligence Index vs. Output Tokens per TaskIntelligence Index Token UseIntelligence Index vs. Token Use

### Output Tokens per Intelligence Index
theanonymousone10860
🟠 redditPlot twist: Gemini 4 Argon tops Val AI benchmark on speed, cost and accuracy!
ClaudeAI
software-boulder9651
🟠 redditGemini 4 Argon Releases
artificial
kairosdev41
🟠 redditGoogle cooked OpenAI and Anthropic with Gemini 4 Argon
artificial
DataRemarkable709315585
🟠 redditGoogle Gemini 4 scores same as GPT 6 Astra on Artificial Analysis Benchmark, while costing 40% less.
singularity
Conscious_Warrior22355
🟠 redditGemini 4 - High Inteligence Index (53) and low on cost (1/3 of Opus 5.5)
singularity
Sea_Physics4019939
🟠 redditGoogle logged claude and openai
singularity
Independent-Wind4462240
🟠 redditGemini 4 being private like Mythos
singularity
usualuzi4812
🟠 redditGemini 4 Argon Benchmarks
singularity
Every_Foundation519713120
🟠 redditGoogle’s unreleased Gemini 4 Argon may have just leaked—and it tops 12 of 18 benchmarks against Fable 5.1, Opus 5.5 and GPT-6 Astra, including 19.6% vs GPT-6 Astra’s 5.4% on autonomous legal work
singularity
141_133722896
🟧 echo.blog ⭐Announces Gemini 4 Argon, 'our next era of frontier intelligence': frontier performance across real-world software engineering, enterprise kGoogle (Koray Kavukcuoglu, SVP Google DeepMind and Chief AI Architect)——
🟠 redditGoogle Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead On the Artificial Analysis Intelligence Index v4.3.2
OpenAI
balianone1128
🟠 redditGemini 4 Argon solved hallucinations.
singularity
drhenriquesoares1519281
🟠 redditGemini-4-argon debuts at 1st on arenai.ai's text arena, and 8th on webdev
singularity
DeArgonaut465
🟧 hnGoogle Grapples with Employee Skepticism About New Gemini Modelmerksittich131
🟠 redditGemini 4 Argon is benchmaxxed...
singularity
PrisonOfH0pe015
🟠 redditOpenAI halved its own pricing from GPT-6 Astra to match Google gemini argon
OpenAI
Domingues_tech34262
🟠 redditReview Gemini 4 - disappointment
singularity
Admirable-Cell-265865
🟠 redditWhile not at the top for coding, Gemini 4 does well on other AI Productivity Indexes
singularity
Marimo1887126
🟠 redditGoogle deepmind engineer denied bloomberg report
singularity
Independent-Wind446219145
🟠 redditGoogle has more powerful model than argon internally
singularity
Independent-Wind446250689
🟠 redditWhile Claude and GPT are still the two best choices, Gemini seems to be catching up on coding agent index with agy-cli
singularity
Marimo1885810
🟠 redditGoogle’s new Gemini 4 Argon is already hanging with Claude’s best on 3D game generation
singularity
141_133716033

Interpretation history

Decision trace