Google DeepMind announced Gemini 4 Argon on 2026-09-30 as a frontier flagship claiming state-of-the-art real-work results (77.9% DeepSWE v1.1, leads on the Vals Index and CWE-bench), 1M-token input and output, and $2/$10 per-million 'introductory' pricing — with phased access that gates cyber capability to vetted defenders via 'Fairwind' under US-government pre-release evaluation rather than opening general availability. The supplied snippets do not directly document the announcement itself: the closest hit (Dataconomy, Sept 21) reports the pre-release leak — an Argon Arena listing near 88% DeepSWE at suspected $2.25/$11.25 pricing — consistent with the case's recorded pre-ship figures (88.7%/2M context) and its leak-vs-ship downgrade to 77.9%/1M. The counter-picture that Argon is benchmaxxed — Artificial Analysis at #8/223, arena-style reviews ranking it below Fable 5/Opus 5.5/GPT-6 Sol, and Bloomberg-reported internal doubts it struggles with real coding work — rests on the case's own evidence titles, not on the supplied web hits. The snippets do corroborate the surrounding market: Google's Sept 2 Gemini 3.8 Flash launch already set the intro-price-then-raise pattern ($0.75/$3.75 doubling on Jan 1, 2027) and a gated 'Flash Cyber' variant, rival specs (GPT-6 Astra 74.1% DeepSWE, Opus 5.5 leading agentic coding) frame exactly the displacement question the case poses, and one comparison site independently voices the case's core lens — '13x more per output token for a lead of twelve points.'
| source | object | author | score | comments |
| 🟧 hn | Gemini 4 ArgonRetrieved article excerptOpen article · Retrieved 2026-09-30T21:38:10.861909+00:00 # Gemini 4 Argon: our next era of frontier intelligence
Sep 30, 2026
|
- [x.com](https://twitter.com/intent/tweet?text=Gemini%204%20Argon%3A%20our%20next%20era%20of%20frontier%20intelligence%20%40google&url=https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/)
- [Facebook](https://www.facebook.com/sharer/sharer.php?caption=Gemini%204%20Argon%3A%20our%20next%20era%20of%20frontier%20intelligence&u=https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/)
- [LinkedIn](https://www.linkedin.com/shareArticle?mini=true&url=https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/&title=Gemini%204%20Argon%3A%20our%20next%20era%20of%20frontier%20intelligence)
- Mail
- Copy link
Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.
---
[Koray Kavukcuoglu
SVP, Google DeepMind and Chief AI Architect, Google](https://blog.google/authors/koray-kavukcuoglu/)
Share
- [x.com](https://twitter.com/intent/tweet?text=Gemini%204%20Argon%3A%20our%20next%20era%20of%20frontier%20intelligence%20%40google&url=https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/)
- [Facebook](https://www.facebook.com/sharer/sharer.php?caption=Gemini%204%20Argon%3A%20our%20next%20era%20of%20frontier%20intelligence&u=https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/)
- [LinkedIn](https://www.linkedin.com/shareArticle?mini=true&url=https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/&title=Gemini%204%20Argon%3A%20our%20next%20era%20of%20frontier%20intelligence)
- Mail
- Copy link
---
Stylized promotional blog key art graphic with modern editorial branding and the text "Gemini 4 Argon"
Your browser does not support the audio element.
Listen to article
[[duration]] minutes
This content is generated by Google AI. Generative AI is experimental
Voice
Speed
Voice
Speed
0.75X
1X
1.5X
2X
Read AI-generated summary
- Google’s new Gemini 4 Argon model brings advanced reasoning to complex, long-horizon professional tasks.
- The model features an industry-leading 1 million token limit for deep, multi-step problem solving.
- It excels at coding, financial research, legal drafting, and autonomous cybersecurity vulnerability patching.
- Argon is currently rolling out to trusted cyber defenders through the Fairwind Program.
- Google is prioritizing safety and rigorous testing before a wider release to the public.
Summaries were generated by Google AI. Generative AI is experimental.
In this article
---
Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our [Fairwind Program](https://deepmind.google/fairwind-program/). Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Google. It delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.
Safely releasing frontier capabilities at this level requires a phased approach. We are actively engaged in the U.S. government’s voluntary process for pre-release model access while we gradually expand access. We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.
Argon will launch at an introductory price
[1](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/#footnote-1)
of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
## Changing how we work and build at Google
Gemini 4 Argon is already powering our internal workflows, with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality. It’s helping teams build faster and push the boundaries of engineering productivity and accelerating breakthroughs:
- **Quantum algorithmic optimization:** Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.
- **Memory efficiency:** A team of Argon agents analyzed fleet-wide profiling telemetry to autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
- **Large Scale Codebase Migrations and Optimizations:** Argon agents are working on migrating C/C++ codebases to Rust across Google—scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel. Given the criticality of many of these systems, such large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production.
For libgav1, Google's open source software for decoding video, Argon agents took an existing Rust port and replaced 32K lines of SIMD code by running many rounds of profile-guided experiments, studying the compiler's output, producing safe Rust so the compiler would vectorize it automatically. The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.
## Working harder on your most complex problems
To support Gemini 4 Argon’s capabilities across longer, more complex use cases, we are significantly expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens. When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go.
a benchmark chart showing Gemini 4 Argon capabilities
## Enabling coding and enterprise workflows across domains
Gemini 4 Argon’s capabilities across coding, reasoning, and multimodality and its ability to sustain long, multi-step tasks enable it to excel across a range of enterprise workflows.
Google engineers have been using Argon for their daily tasks, from everyday debugging to large-scale codebase migrations and algorithm designs. It sets a new state of the art on DeepSWE v1.1 (77.9%), which measures a model’s performance in real-world long-horizon software engineering tasks.
Beyond coding, Argon is the leading model on the [Vals Index](https://www.vals.ai/benchmarks/vals_index), which measures economic impact across finance, coding, legal, and tax work, with every sector weighted by its contribution to U.S. GDP. We see similarly leading performance across other domain specific evaluations, like Vals Finance Agent v2 (multi-step financial research) and Harvey’s Legal Agent Benchmark (legal research and drafting). On AutomationBench, Zapier’s benchmark measuring end-to-end execution across core business functions, Argon ranks #1 with a score of 51.3%.
Argon is also uniquely strong when knowledge work requires visual understanding. It’s able to drive professional chart analysis, identify details from long videos, and take action based on a series of documents. For example, on LVBench, which measures long video understanding, Argon is state of the art with a score of 91.7%.
DeepSwe evaluation chart
Vals index chart
Val's finance benchmark chart
Harveys benchmark chart
automation bench chart
## Leading in defensive cybersecurity
To better equip cyber defenders for the new era of cyberattacks, we trained Gemini 4 Argon to be highly capable at cybersecurity defense. Argon can autonomously find, validate, and patch critical software vulnerabilities. For trusted defenders and our own internal teams at Google, we’ll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities.
[Wiz](https://www.wiz.io/) is already using Argon for cybersecurity defense through its [Scan for Good](https://www.wiz.io/scan-for-good) initiative – a program dedicated to protecting critical public infrastructure for free by finding and remediating high-risk exposures. In an early demonstration of its impact, the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, identifying a severe risk that previous frontier models had missed.
On [CWE-bench v1](https://cwe-bench.com/), which evaluates the model’s ability to remediate security vulnerabilities, Argon ties for first place with a top score of 68%, building on 3.8 Flash Cyber’s frontier performance on [CWE-bench v0](https://cwe-bench.com/?v=v0).
CWE benchmark chart
Gemini 4 Argon demonstrates impressive leaps in vulnerability discovery over 3.8 Flash Cyber. For example:
- On Google’s internal comprehensive vulnerability benchmark, Argon uncovered a wide range of exposures across complex codebases spanning 20 programming languages.
- On Wiz’s internal black-box penetration testing benchmark, which tests a model’s ability to analyze live web systems without source code, Argon outperforms 3.8 Flash Cyber in discovering the attack surface, identifying vulnerabilities, and producing proof-of-concept evidence to validate them.
security vulnerabilities chart
## Strengthening frontier safeguards before broad availability
Before rolling out Gemini 4 Argon broadly, we’re continuing to strengthen critical frontier safeguards across four main areas:
**Defending against misuse:** To prevent bad actors from using Argon for cyber or chemical, biological, radiological, and nuclear (CBRN) attacks, the model is designed to refuse harmful requests while preserving legitimate, dual-use scientific research, as per our [Frontier Safety Framework](https://deepmind.google/blog/strengthening-our-frontier-safety-framework/). We are strengthening the robustness of our safeguards for this launch, including improving our techniques to monitor the model’s [internal activations](https://arxiv.org/abs/2601.11516) to spot misuse. These safeguards underwent robustness testing by internal and external red teams using a combination of manual and automated attack methods.
**Defending against prompt injection attacks:** Argon is also our most resilient model yet against indirect prompt injections, where malicious instructions or context are used to hijack a model’s behavior. These are complex attacks that require constant vigilance and multiple layers of defense. Through automated red teaming and adversarial training, Gemini 4 Argon is leading in prompt injection robustness on the Gray Swan’s Indirect Prompt Injection (IPI) benchmark.
Gray Swan evaluation
**Monitoring for misalignment:** In order to prevent Argon from stepping out of bounds to try to accomplish a task in a way that goes beyond the user’s intentions, we are deploying misalignment mitigations that monitor Argon’s chain-of-thought and actions and stop execution when necessary.
We used a similar system to monitor our training runs and send alerts to a dedicated incident response team, taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We [strongly encourage the rest of the industry](https://institute.deepmind.com/essays/the-case-for-reasoning-transparency/) to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.
**Hardeni | bradleyg223 | 1704 | 1188 |
| 🟠 reddit | Gemini 4 is out: The competition has woken up ClaudeAI Retrieved article excerptOpen article · Retrieved 2026-09-30T21:38:08.844028+00:00 # Prove your humanity
We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.
[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)
[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us) | monsieurcliffe | 1198 | 278 |
| 🟠 reddit | Wake up babe, it's Gemini 4 fr OpenAI Retrieved article excerptOpen article · Retrieved 2026-09-30T21:38:04.445609+00:00 # Prove your humanity
We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.
[Reddit, Inc. © "2026". All rights reserved.](https://www.redditinc.com/)
[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us) | MrTimeHacker1 | 478 | 87 |
| 🟠 reddit | Gemini 4 Crushes Benchmarks, But Google Employees State The Model Struggles With Real Work singularity | Neurogence | 173 | 90 |
| 🟧 hn | Gemini 4 Argon (High): Intelligence, Performance and Price AnalysisRetrieved article excerptOpen article · Retrieved 2026-09-30T21:38:12.095195+00:00 [Artificial Analysis](https://artificialanalysis.ai/)
K
Google logo
[Google](https://artificialanalysis.ai/models/creators/google)
•
Proprietary model
•
Released September 2026
# Gemini 4 Argon (High) Intelligence, Performance & Price Analysis
Compare[Try it out](https://artificialanalysis.ai/microevals) [API Provider Benchmarks](https://artificialanalysis.ai/models/gemini-4-argon/providers)
### Model summary
#### [Intelligence](https://artificialanalysis.ai/models/gemini-4-argon#intelligence)Updated
#8 / 223
53
Artificial Analysis Intelligence Index
4 out of 4 units for Intelligence.
#### Speed
N/A
Output tokens per second
Unknown out of 4 units for Speed.
#### [Cost](https://artificialanalysis.ai/models/gemini-4-argon#price-cost)
#77 / 223
In $2.00Out $10.00Cache Discount 95%
$1.99
Cost per Intelligence Index task
3 out of 4 units for Cost.
#### [Verbosity](https://artificialanalysis.ai/models/gemini-4-argon#token-use)
#74 / 223
110M
Output tokens from Intelligence Index
3 out of 4 units for Verbosity.
### Comparison Summary
Gemini 4 Argon (High) is amongst the leading models in intelligence and reasonably priced when comparing to other models of similar price. The model supports text and image input, outputs text, and has a 1M tokens context window.
Gemini 4 Argon (High) scores 53 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median: 26). When evaluating the Intelligence Index, it generated 110M tokens, which is somewhat verbose in comparison to the median of 82M.
Pricing for Gemini 4 Argon (High) is $2.00 per 1M input tokens (moderately priced, median: $2.00) and $10.00 per 1M output tokens (moderately priced, median: $10.00). On average, it costs $1.99 per task to evaluate Gemini 4 Argon (High) on the Intelligence Index.
### Technical specifications
| | |
| --- | --- |
| Reasoning | Yes This page shows the reasoning version of this model. A non-reasoning variant may also exist. |
| Input modality | Supports: text and image |
| Output modality | Supports: text |
| Context window | 1M ~1500 A4 pages of size 12 Arial font |
### 223 models in this class
Metrics are compared against models of the same class:
- Non-reasoning models → compared only with other non-reasoning models
- Reasoning models → compared across both reasoning and non-reasoning
- Open weights models → compared only with other open weights models of the same size class:
- Tiny: ≤4B parameters
- Small: 4B–40B parameters
- Medium: 40B–150B parameters
- Large: >150B parameters
- Proprietary models → compared across proprietary and open weights models of the same price range, using a blended 3:1 input/output price ratio:
- <$0.15 per 1M tokens
- $0.15–$1 per 1M tokens
- >$1 per 1M tokens
Highlights
Updated
### [Intelligence](https://artificialanalysis.ai/models/gemini-4-argon#intelligence)
Artificial Analysis Intelligence Index · Higher is better
Not publicly available
### Speed
Output tokens per second · Higher is better
### [Cost per Task](https://artificialanalysis.ai/models/gemini-4-argon#price-cost)
Weighted average cost (USD) per Intelligence Index task · Lower is better
Not publicly available
Prompt Options
## IntelligenceUpdated
### [Artificial Analysis Intelligence Index](https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index)
Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
25 of 687 models
Add model from specific provider
Not publicly available
### Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See [Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) for further details, including a breakdown of each evaluation and how we run them.
Open Weights / ProprietaryReasoning / Non-ReasoningText Only / Multimodal Inputs
### Artificial Analysis Intelligence Index by Open Weights / Proprietary
Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
25 of 687 models
Add model from specific provider
Not publicly available
ProprietaryOpen WeightsOpen Weights (Commercial Use Restricted)
### Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See [Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) for further details, including a breakdown of each evaluation and how we run them.
### Open Weights
Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if commercial use is limited by conditions, and as 'Non-commercial' if the license prohibits commercial use.
## [Capability Indexes](https://artificialanalysis.ai/models/capabilities)
Measures the performance of models on specific capabilities and industries
Finance & AccountingStrategy & OpsLegalEngineeringEconomics
### [Artificial Analysis Finance & Accounting Index](https://artificialanalysis.ai/models/capabilities/finance-and-accounting)
Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better
25 of 194 models
Add model from specific provider
Not publicly available
## [Benchmarks](https://artificialanalysis.ai/evaluations)
### Intelligence Evaluations
Intelligence evaluations measured independently by Artificial Analysis · Higher is better
CodingAgenticTool UsePrivate DatasetUser InteractionFinanceMedicalLegalIntelligence IndexLong ContextMultimodalInstruction FollowingFaithfulnessWritingBusiness[See more](https://artificialanalysis.ai/evaluations)
18 of 27 evaluations
25 of 687 models
Add model from specific provider
Not publicly available
[AA-Briefcase v1.1](https://artificialanalysis.ai/evaluations/aa-briefcase)Updated
Agentic knowledge work, (Elo-500)/2000
[GDPval-AA v2.1](https://artificialanalysis.ai/evaluations/gdpval-aa)Updated
Agentic real-world work tasks, (Elo-500)/2000
[AutomationBench-AA](https://artificialanalysis.ai/evaluations/automationbench-aa)Updated
Agentic SaaS workflows
[Terminal-Bench 4.0](https://artificialanalysis.ai/evaluations/terminalbench-4-0)New
Agentic coding & terminal use
[SciCode](https://artificialanalysis.ai/evaluations/scicode)Under review
Coding
[Humanity's Last Exam](https://artificialanalysis.ai/evaluations/humanitys-last-exam)
Reasoning & knowledge
[GDP.pdf](https://artificialanalysis.ai/evaluations/gdp-pdf)New
Professional document reasoning, All-pass
[CritPt](https://artificialanalysis.ai/evaluations/critpt)Under review
Physics reasoning
[AA-Omniscience Accuracy](https://artificialanalysis.ai/evaluations/omniscience)
Knowledge
[AA-Omniscience Non-Hallucination Rate](https://artificialanalysis.ai/evaluations/omniscience)
1 - hallucination rate
[AA-LCR v1.1](https://artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning)
Long context reasoning
[Harvey LAB-AA](https://artificialanalysis.ai/evaluations/harvey-lab-aa)
Legal agentic work, criterion pass rate
[EnterpriseOps-Gym-AA](https://artificialanalysis.ai/evaluations/enterprise-ops-gym-aa)
Agentic business operations
[Terminal-Bench-Science 0.1](https://artificialanalysis.ai/evaluations/terminal-bench-science)New
Agentic scientific research workflows in a terminal
[AA-AnalystAgent](https://artificialanalysis.ai/evaluations/aa-analyst-agent)
Quantitative analysis on spreadsheets & documents
[ITBench-AA](https://artificialanalysis.ai/evaluations/itbench-aa)
Kubernetes incident root-cause analysis
[MMMU-Pro](https://artificialanalysis.ai/evaluations/mmmu-pro)
Visual reasoning
[MLCR-AA](https://artificialanalysis.ai/evaluations/mlcr-aa)
Medical long context reasoning
### Intelligence Evaluation Relevance
While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.
### Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See [Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) for further details, including a breakdown of each evaluation and how we run them.
### AA-Briefcase v1.1Updated
AA-Briefcase EloAA-Briefcase Rubric Score (%)Analytical Quality & Presentation Elo
### AA-Briefcase Elo
AA-Briefcase v1.1 is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better
25 of 211 models
Add model from specific provider
Not publicly available
### AA-Briefcase Elo
AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0.
### AA-Omniscience
AA-Omniscience IndexAA-Omniscience AccuracyAA-Omniscience Hallucination Rate
### AA-Omniscience Index
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.
25 of 562 models
Add model from specific provider
Not publicly available
### AA-Omniscience Index
AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.
## Intelligence Index Comparisons
Intelligence Index vs. Cost per TaskIntelligence Index vs. Time per TaskIntelligence Index vs. Output SpeedIntelligence Index vs. End-to-End Response Time
### Intelligence Index vs. Cost per Intelligence Index Task
Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
25 of 687 models
Most attractive quadrant
Pareto line
GoogleXiaomiOpenAIAnthropicSpaceXAIMetaZ AIDeepSeek
### Cost per Intelligence Index Task
Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.
### Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See [Intelligence Index methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking) for further details, including a breakdown of each evaluation and how we run them.
## Token Use
Output Tokens per TaskIntelligence Index vs. Output Tokens per TaskIntelligence Index Token UseIntelligence Index vs. Token Use
### Output Tokens per Intelligence Index | theanonymousone | 108 | 60 |
| 🟠 reddit | Plot twist: Gemini 4 Argon tops Val AI benchmark on speed, cost and accuracy! ClaudeAI | software-boulder | 96 | 51 |
| 🟠 reddit | Gemini 4 Argon Releases artificial | kairosdev | 4 | 1 |
| 🟠 reddit | Google cooked OpenAI and Anthropic with Gemini 4 Argon artificial | DataRemarkable7093 | 155 | 85 |
| 🟠 reddit | Google Gemini 4 scores same as GPT 6 Astra on Artificial Analysis Benchmark, while costing 40% less. singularity | Conscious_Warrior | 223 | 55 |
| 🟠 reddit | Gemini 4 - High Inteligence Index (53) and low on cost (1/3 of Opus 5.5) singularity | Sea_Physics401 | 99 | 39 |
| 🟠 reddit | Google logged claude and openai singularity | Independent-Wind4462 | 2 | 40 |
| 🟠 reddit | Gemini 4 being private like Mythos singularity | usualuzi | 48 | 12 |
| 🟠 reddit | Gemini 4 Argon Benchmarks singularity | Every_Foundation5197 | 131 | 20 |
| 🟠 reddit | Google’s unreleased Gemini 4 Argon may have just leaked—and it tops 12 of 18 benchmarks against Fable 5.1, Opus 5.5 and GPT-6 Astra, including 19.6% vs GPT-6 Astra’s 5.4% on autonomous legal work singularity | 141_1337 | 228 | 96 |
| 🟧 echo.blog ⭐ | Announces Gemini 4 Argon, 'our next era of frontier intelligence': frontier performance across real-world software engineering, enterprise k | Google (Koray Kavukcuoglu, SVP Google DeepMind and Chief AI Architect) | — | — |
| 🟠 reddit | Google Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead On the Artificial Analysis Intelligence Index v4.3.2 OpenAI | balianone | 11 | 28 |
| 🟠 reddit | Gemini 4 Argon solved hallucinations. singularity | drhenriquesoares | 1519 | 281 |
| 🟠 reddit | Gemini-4-argon debuts at 1st on arenai.ai's text arena, and 8th on webdev singularity | DeArgonaut | 46 | 5 |
| 🟧 hn | Google Grapples with Employee Skepticism About New Gemini Model | merksittich | 13 | 1 |
| 🟠 reddit | Gemini 4 Argon is benchmaxxed... singularity | PrisonOfH0pe | 0 | 15 |
| 🟠 reddit | OpenAI halved its own pricing from GPT-6 Astra to match Google gemini argon OpenAI | Domingues_tech | 342 | 62 |
| 🟠 reddit | Review Gemini 4 - disappointment singularity | Admirable-Cell-2658 | 6 | 5 |
| 🟠 reddit | While not at the top for coding, Gemini 4 does well on other AI Productivity Indexes singularity | Marimo188 | 71 | 26 |
| 🟠 reddit | Google deepmind engineer denied bloomberg report singularity | Independent-Wind4462 | 191 | 45 |
| 🟠 reddit | Google has more powerful model than argon internally singularity | Independent-Wind4462 | 506 | 89 |
| 🟠 reddit | While Claude and GPT are still the two best choices, Gemini seems to be catching up on coding agent index with agy-cli singularity | Marimo188 | 58 | 10 |
| 🟠 reddit | Google’s new Gemini 4 Argon is already hanging with Claude’s best on 3D game generation singularity | 141_1337 | 160 | 33 |
2026-10-04T23:29:37Z
No meaning change: the velocity spike is the lone hands-on 3D-game post (~160 pts) crossing its cohort p90 against a ~0 pts/h reception base — an engagement increment on an already-priced single-source item — and the magnitude-valve reading is still the announcement wave, structurally frozen while the Fairwind gate blocks any new communities, implementations, or evaluator lines. The case remains a gated, price-led knowledge-work challenger that is not a coding-agent default, waiting on its external re-triggers.
2026-10-04T00:43:14Z
grounded: converges/high — Converges on two of Scott's core positions at once, with consequential parties: Bloomberg-sourced Google-employee admissions that Argon 'crushes benchmarks but
2026-10-04T00:34:24Z
No new substance: the lone positive hands-on item (community 3D-game-generation test) grew ~15→149 pts but remains a single source with unclear provenance while the access gate keeps independent hands-on evidence impossible, and everything else is small engagement increments on already-priced announcement items — the velocity-spike sensor flags are noise against a reception base now at ~0 pts/h. Meaning is unchanged: a gated, price-led knowledge-work challenger that is not a coding-agent default, waiting on its external re-triggers.
2026-10-02T16:53:14Z
The only new substantive item — a community 3D-game-generation hands-on test (15 pts, 3 comments) — is the first faint positive hands-on signal, but too narrow, weak and provenance-unclear (the model is gated) to dent the five-evaluator split or count as the 'first hands-on agent reports' re-trigger. Reception is now ~1.6% of peak (~25 pts/h vs ~1576, steady, 93rd pct only against its own age cohort); the magnitude-valve cross-platform spread is the already-priced announcement wave, not expanding periphery — no new communities, implementations, or evaluator lines are possible while the gate holds — so low heat holds and the case waits on its external re-triggers.
2026-10-02T16:26:23Z
evidence attached: reddit.post.1wvxh40 — Hands-on community test showing Argon competitive on 3D game generation is a (weak, single-source) positive capability signal against the case's recorded counter-signals.
2026-10-02T08:24:47Z
The AA Coding Agent Index adds the case's first model+harness measurement: Argon in Google's own Antigravity CLI scores 64 — under Sonnet 5.5's 68 and level with Sol 6.1's 63 at ~5.6x Sol's per-task cost — so the intro-pricing cost advantage does not survive agentic-coding task economics, consolidating 'not a coding-agent default' and moving the displacement question wholly onto knowledge-work cost pressure. With rate at ~27 pts/h (~2% of peak, cooling, 66h old) and the remaining motion being derivative amplification (the internal-model post crossing 400 pts on the same checkpoint narrative) while the gate keeps hands-on evidence impossible, heat drops to low: the magnitude-valve spread reflects the already-priced announcement wave, not an expanding periphery — no new platforms, communities, or evaluator motion beyond that one benchmark post.
2026-10-02T08:23:07Z
evidence attached: reddit.post.1wvnh2d — Fresh Artificial Analysis Coding Agent Index numbers (Argon 64 at $5.84 vs Sol 63 at $1.04, Sonnet 5.5 top at $14.19) directly bear on Argon's default-workload question and also contextualise the Sol cost-efficiency and Sonnet 5.5 token-economics cases.
2026-10-01T22:15:21Z
The unsourced 'Google holds a stronger internal model' post (1wv7jfq, ~160 pts) is derivative, not new substance: its own comment section reads it as the standard hold-back-a-checkpoint pattern, so it folds into the already-tracked leak-vs-ship downgrade (88.7→77.9 DeepSWE, 2M→1M context) and consolidates the checkpoint-release reading without moving the four-evaluator split profile. Meaning is unchanged; temperature keeps falling — rate now ~108 pts/h (~7% of peak, cooling, 98th percentile), no new platforms, evaluator lines, or implementations possible while the gate holds — so medium stays priced on the live watch items, chiefly verification or kill of the still-uncorroborated Astra price cut.
2026-10-01T20:35:48Z
evidence attached: reddit.post.1wv7jfq — Widely-engaged claim that Google holds a stronger unreleased internal model contextualizes Argon's positioning and adds to the case's tracked counter-signals.
2026-10-01T15:24:24Z
No meaning change: the DeepMind engineer's public denial of the Bloomberg report is an interested-party counter-signal the split profile already anticipated, and the velocity spike is the community amplifying the still-unverified Astra price-cut screenshot (303 pts, no pricing-page or independent confirmation) rather than periphery expansion — evidence set is unchanged and momentum sits at ~14% of peak. What narrows is the live edge: the case now waits on verification or kill of the Astra price response, the first re-trigger that would materialize the cost-pressure thesis, and medium heat holds on that pending verification, not on engagement.
2026-10-01T14:32:27Z
evidence attached: reddit.post.1wuz9sg — A DeepMind engineer publicly denying the Bloomberg internal-doubts report is the direct counter-signal dispute the Argon case is already tracking, even if interested-party.
2026-10-01T10:41:49Z
The Mercor APEX attachment completes a four-evaluator independent profile that sharpens the displacement hypothesis into a split: Argon looks genuinely competitive on knowledge-work/productivity indexes (Vals lead, APEX strength, ~15% hallucination rate) while AA, two arena-style reviews and Bloomberg-reported internal doubts agree it is not a coding-work leader — so the live question narrows to cost-led pressure on Astra's knowledge-work turf, with Opus 5.5's coding-agent default increasingly out of reach until hands-on access exists. Measured heat (99.6th percentile, magnitude-valve spread) overstates periphery: rate is ~17% of peak, cooling, and new arrivals are single-digit derivative posts while the gate blocks any new implementations — medium holds on the live watch items (gate opening, confirmed Astra price response, Fable response), not on engagement.
2026-10-01T10:23:58Z
evidence attached: reddit.post.1wuuu14 — Third-party Mercor APEX leaderboard spread showing Argon strong on productivity indexes despite the coding gap is independent benchmark context the Argon re-judgment must weigh.
2026-10-01T07:44:57Z
The Arena-style review ranking Argon below Fable 5, Opus 5, and GPT-6 Sol is a third independent line converging on the already-settled 'price-led competitor, not capability leader' read — convergence, not a shift in meaning. The velocity spike is a re-surge of an initial-wave post while aggregate momentum cools and new arrivals are single-digit derivative posts, so the magnitude-valve spread reading reflects the wave's thinning tail rather than expanding periphery; the case holds on the Fairwind gate and stays at medium heat.
2026-10-01T07:24:07Z
evidence attached: reddit.post.1wuswuc — Independent arena-style evaluation ranking Gemini 4 Argon below Fable 5, Opus 5, and GPT-6 Sol — another counter-signal the case explicitly tracks against Google's frontier-leadership claim.
2026-10-01T06:29:39Z
The OpenAI Astra price-halving report folds in as an unconfirmed watch item, not a shift: the case's meaning stays 'price-led competitor, not capability leader,' and with the reception wave now visibly cresting (current rate ~20% of peak, momentum flipped to cooling, newest arrivals single-digit-score posts) the case enters a holding pattern on the Fairwind gate. Heat steps down from high to medium despite the magnitude-valve spread reading — aggregate engagement is still top-percentile but the derivative stream is thinning and nothing decision-relevant can arrive until access opens; the gate opening, first hands-on agent reports, or confirmed competitor pricing will re-trigger loudly.
2026-10-01T06:23:30Z
evidence attached: reddit.post.1wurlh9 — Reported OpenAI halving of Astra pricing to match Argon is direct displacement-pressure evidence for the case, though a low-traffic post that needs confirmation.
2026-10-01T04:51:36Z
The two post-corroboration attachments are convergence, not new substance: the Bloomberg HN echo and a low-traction benchmaxx thread (whose one fresh detail — an early vibe-check placing Argon on the Pareto frontier mainly on price vs GPT-6.1 Sol — confirms the existing read) restate the already-established counter-picture. The case's meaning is settled as 'price-led competitor, not capability leader,' and it is now purely waiting on the Fairwind gate before the displacement question can be tested; heat stays high on the spread reading alone — 99.7th-percentile velocity, three platforms, top-decile magnitude, momentum still accelerating — because the reception wave keeps spawning derivative coverage faster than it decays.
2026-10-01T02:29:13Z
evidence attached: reddit.post.1wumtj3 — Early hands-on counter-signal (benchmaxxing, Terminal-Bench doubts) plus a Pareto-pricing counterpoint directly bearing on whether Argon holds up against Sol/Astra as an agent default.
2026-10-01T02:29:13Z
evidence attached: hn.story.49913950 — The Bloomberg employee-skepticism report on the new Gemini model surfacing on HN is the internal-doubt counter-signal the Argon case's resolution hinges on.
2026-10-01T01:00:54Z
The case moves from 'one announcement echoed' to a corroborated counter-picture: two independent measurement lines (AA v4.3.2 #8/223; arena #1 text but #8 webdev) plus Bloomberg internal-doubt reporting now converge on 'price-competitive, not real-work frontier-leading,' and AA's ~15% hallucination rate is the first genuinely differentiated positive signal. The displacement question stays untestable while the Fairwind gate holds; heat stays high because the periphery is still expanding (arena/AA derivative posts, Bloomberg follow-on) at 99th-percentile velocity despite the rate cooling from peak.
2026-10-01T00:29:45Z
evidence attached: reddit.post.1wuicar — Independent arena data — #1 text arena but #8 webdev — directly feeds the case's benchmark-claims-vs-real-work tension.
2026-10-01T00:29:45Z
evidence attached: reddit.post.1wuj72j — Viral (234pt, 93%) reception claim that Argon 'solved hallucinations' is spread evidence for the open Argon reception case.
2026-10-01T00:29:45Z
evidence attached: reddit.post.1wuiwbq — An Artificial Analysis v4.3.2 reading showing Argon closing but not taking the lead feeds the case's open question of whether Argon's frontier standing holds.
2026-09-30T21:52:31Z
grounded: converges/high — Converges: Bloomberg-sourced reporting that Google employees say Argon 'crushes benchmarks but struggles with real work' is a consequential party — the vendor's
2026-09-30T21:42:03Z
case created — One Google announcement echoed across HN (522/291), r/ClaudeAI, r/singularity and r/artificial with same-day AA/Vals measurements and Bloomberg internal-doubt reporting — ten proposals are facets of a single release episode and consolidate here.