Grok is SpaceXAI/xAI’s model family for coding, agentic tasks, and knowledge work; the supplied official snippets document Grok 4.5’s launch and describe Grok 4.6 as focused on long-running agents. They do not establish the case’s claimed Grok 4.7 release: third-party snippets describe 4.7 as forthcoming or unreleased, with unpublished pricing and conflicting specifications. The official Grok 4.5 announcement lists $2 per million input tokens and $6 per million output tokens and claims superior token efficiency; reporting on Grok 4.6 gives the same starting rates but says they double for prompts reaching 200,000 tokens. These snippets therefore support an existing agent-model price/performance pitch, not verified Grok 4.7 availability, unchanged pricing, or frontier-comparable multi-hour performance.
The supported price/performance pitch adds no established position beyond Scott’s AI Unit Economics and Model-Plus-Harness Benchmark Unit: compare cost per successful task under a disclosed harness, not token rates alone; no supplied radar page already tracks this exact release. This warrants verification because Ask terminal agent already uses a `grok → smart → cheap` alternate routing chain, but the grounding does not establish Grok 4.7 availability, pricing or performance, so it supplies neither a justified model switch nor credible convergence with his long-running-agent architecture.
ip:concept.ai-unit-economicsip:concept.model-plus-harness-benchmark-unitdev:project.askradar:hidden-reasoning-real-task-costsradar:concept.long-running-agentsradar:concept.coding-agent-evaluation
queries asked of Scott's wikis
- coding agent harness model selection and switching
- long-running autonomous agents reliability evaluation
- cost per completed task versus token pricing
- context compaction caching long-context inference costs
- knowledge-work agents repository-specific benchmarks
now 0 pts/hpeak 13 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 506h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
| source | object | author | score | comments |
| 🟠 reddit | Introducing Grok 4.7 singularity Retrieved article excerptOpen article · Retrieved 2026-09-21T16:23:38.283904+00:00 [Back to news](https://x.ai/news)Sep 21, 2026
# Introducing Grok 4.7
SpaceXAI's most powerful model for coding and knowledge work. Twice as fast, at half the price of comparable models.
[Try for free](https://x.ai/build?utm_source=website&utm_medium=referral&utm_campaign=grok-4-7-blog&utm_content=build-cta)[Start building](https://console.x.ai?utm_source=website&utm_medium=referral&utm_campaign=grok-4-7-blog&utm_content=build-cta)
Grok 4.7 is our most capable model for coding and knowledge work. It works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date. Served at the same price and speed as [Grok 4.6](https://x.ai/news/grok-4-6), it is highly competitive in its class.
A scatter and line chart comparing Fable 5.1, Opus 5, Grok 4.7, GPT-5.6 Sol, and Sonnet 5 scores against average cost per task.55%CursorBench 4.0 score50%45%40%35%30%25%20%$18$15$12$9$6$3$0Average cost per taskFable 5.1Opus 5GPT-5.6 SolSonnet 5Grok 4.7
CostTokensSteps
On CursorBench 4.0, which stresses longer-running coding tasks, Grok 4.7 is at the frontier in price-performance.
## [Model Improvements](https://x.ai/news/grok-4-7#model-improvements)
Grok 4.7 uses a new, larger base model compared to [Grok 4.6](https://x.ai/news/grok-4-6). It was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. The model is better at verifying its own work and managing longer context. We also trained Grok 4.7 to natively understand the [Grok Bot](https://x.ai/bot) harness, making it better at conversational tasks and general knowledge work.
Grok 4.7 xHigh
Grok 4.6 High
GPT-5.6 Sol Max
Fable 5.1 Max
Input token price, $ per million
$2
$2
$4
$10
Output token price, $ per million
$6
$6
$20
$50
Software engineeringCursorBench 4.0
46.3%
40.4%
41.7%
51.8%
Software engineeringDeepSWE v1.1
71.0%\*
65.2%
72.7%
70.0%
Electrical engineeringEEBench
64.0%
53.0%
39.4%
56.4%
Multi-hour office workAA Briefcase v1.1
1,657
1,546
1,487
1,678
Multi-hour terminal workTerminal-Bench 4.0
38.0%
20.3%
37.3%
57.9%
Legal workHarvey Legal Agent Benchmark
19.6%
15.8%
2.5%
6.7%
Clinical reasoningHealthBench Professional
56.7%
48.5%
60.5%
62.1%
\* high effort
Token prices and benchmark scores for Grok 4.7, Grok 4.6, GPT-5.6 Sol, and Fable 5.1. Benchmarks are CursorBench 4.0, DeepSWE v1.1, EEBench, AA Briefcase v1.1, Terminal-Bench 4.0, Harvey Legal Agent Benchmark, and HealthBench Professional. An asterisk on Grok 4.7 DeepSWE marks a high-effort score.
Grok 4.7 is better at creating documents and presentations. In GDPval and AA Briefcase, AI is asked to work on tasks done by professionals such as lawyers, nurses, and financial analysts. Grok 4.7 improves upon Grok 4.6 on both benchmarks and performs comparably to other frontier models.
Professional knowledge workMulti-hour office workElectrical engineering
Professional knowledge work
GDPval
050010001500Elo score1735Fable 5.1 (max)1695Grok 4.7 (xhigh)1605Grok 4.6 (high)1542GPT-6 Astra (max)
GDPval, AA Briefcase, and EEBench scores comparing Grok 4.7 with Grok 4.6, Fable 5.1, and GPT-6 Astra.
## [Safety & Cybersecurity](https://x.ai/news/grok-4-7#safety--cybersecurity)
Grok 4.7 was built with an entirely new safeguard stack. It is the strongest model we’ve tested on refusals and jailbreak resistance. In dual-use domains like cybersecurity and biological work, it leads on both utility for benign tasks and safe refusal on dangerous ones, topping LatchBio’s biosafety benchmark at 62.4%.
Grok 4.7 balances strong cyber defense capabilities with low refusal rates for legitimate use. It shows the highest safety on HackerBench v0.3, our benchmark for risky and malicious cyber tasks, allowing only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. We’ve also started giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defense research.
## [Pricing and availability](https://x.ai/news/grok-4-7#pricing-and-availability)
Grok 4.7 is available today in [Cursor](https://cursor.com) and [Grok Build](https://x.ai/build). It is also available through the [Grok API](https://console.x.ai?campaign=grok-4-7-blog&utm_source=website&utm_medium=referral&utm_campaign=grok-4-7-blog), third-party coding harnesses, and model routers and cloud platforms.
The model is priced starting at $2 per million input tokens and $6 per million output tokens. We also serve a fast variant with twice the output speed at twice the price.
[Console
Create an API key](https://console.x.ai/team/default/api-keys?campaign=grok-4-7-blog&utm_source=website&utm_medium=referral&utm_campaign=grok-4-7-blog)[docs.x.ai
Read the docs](https://docs.x.ai)
### Try it in Grok Build for free
Get started today at [x.ai/build](https://x.ai/build).
`$ curl -fsSL https://x.ai/cli/install.sh | bash` | AMBNNJ | 389 | 113 |
| 🟧 hn | Grok 4.7Retrieved article excerptOpen article · Retrieved 2026-09-21T16:23:39.204306+00:00 [Back to news](https://x.ai/news)Sep 21, 2026
# Introducing Grok 4.7
SpaceXAI's most powerful model for coding and knowledge work. Twice as fast, at half the price of comparable models.
[Try for free](https://x.ai/build?utm_source=website&utm_medium=referral&utm_campaign=grok-4-7-blog&utm_content=build-cta)[Start building](https://console.x.ai?utm_source=website&utm_medium=referral&utm_campaign=grok-4-7-blog&utm_content=build-cta)
Grok 4.7 is our most capable model for coding and knowledge work. It works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date. Served at the same price and speed as [Grok 4.6](https://x.ai/news/grok-4-6), it is highly competitive in its class.
A scatter and line chart comparing Fable 5.1, Opus 5, Grok 4.7, GPT-5.6 Sol, and Sonnet 5 scores against average cost per task.55%CursorBench 4.0 score50%45%40%35%30%25%20%$18$15$12$9$6$3$0Average cost per taskFable 5.1Opus 5GPT-5.6 SolSonnet 5Grok 4.7
CostTokensSteps
On CursorBench 4.0, which stresses longer-running coding tasks, Grok 4.7 is at the frontier in price-performance.
## [Model Improvements](https://x.ai/news/grok-4-7#model-improvements)
Grok 4.7 uses a new, larger base model compared to [Grok 4.6](https://x.ai/news/grok-4-6). It was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. The model is better at verifying its own work and managing longer context. We also trained Grok 4.7 to natively understand the [Grok Bot](https://x.ai/bot) harness, making it better at conversational tasks and general knowledge work.
Grok 4.7 xHigh
Grok 4.6 High
GPT-5.6 Sol Max
Fable 5.1 Max
Input token price, $ per million
$2
$2
$4
$10
Output token price, $ per million
$6
$6
$20
$50
Software engineeringCursorBench 4.0
46.3%
40.4%
41.7%
51.8%
Software engineeringDeepSWE v1.1
71.0%\*
65.2%
72.7%
70.0%
Electrical engineeringEEBench
64.0%
53.0%
39.4%
56.4%
Multi-hour office workAA Briefcase v1.1
1,657
1,546
1,487
1,678
Multi-hour terminal workTerminal-Bench 4.0
38.0%
20.3%
37.3%
57.9%
Legal workHarvey Legal Agent Benchmark
19.6%
15.8%
2.5%
6.7%
Clinical reasoningHealthBench Professional
56.7%
48.5%
60.5%
62.1%
\* high effort
Token prices and benchmark scores for Grok 4.7, Grok 4.6, GPT-5.6 Sol, and Fable 5.1. Benchmarks are CursorBench 4.0, DeepSWE v1.1, EEBench, AA Briefcase v1.1, Terminal-Bench 4.0, Harvey Legal Agent Benchmark, and HealthBench Professional. An asterisk on Grok 4.7 DeepSWE marks a high-effort score.
Grok 4.7 is better at creating documents and presentations. In GDPval and AA Briefcase, AI is asked to work on tasks done by professionals such as lawyers, nurses, and financial analysts. Grok 4.7 improves upon Grok 4.6 on both benchmarks and performs comparably to other frontier models.
Professional knowledge workMulti-hour office workElectrical engineering
Professional knowledge work
GDPval
050010001500Elo score1735Fable 5.1 (max)1695Grok 4.7 (xhigh)1605Grok 4.6 (high)1542GPT-6 Astra (max)
GDPval, AA Briefcase, and EEBench scores comparing Grok 4.7 with Grok 4.6, Fable 5.1, and GPT-6 Astra.
## [Safety & Cybersecurity](https://x.ai/news/grok-4-7#safety--cybersecurity)
Grok 4.7 was built with an entirely new safeguard stack. It is the strongest model we’ve tested on refusals and jailbreak resistance. In dual-use domains like cybersecurity and biological work, it leads on both utility for benign tasks and safe refusal on dangerous ones, topping LatchBio’s biosafety benchmark at 62.4%.
Grok 4.7 balances strong cyber defense capabilities with low refusal rates for legitimate use. It shows the highest safety on HackerBench v0.3, our benchmark for risky and malicious cyber tasks, allowing only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. We’ve also started giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defense research.
## [Pricing and availability](https://x.ai/news/grok-4-7#pricing-and-availability)
Grok 4.7 is available today in [Cursor](https://cursor.com) and [Grok Build](https://x.ai/build). It is also available through the [Grok API](https://console.x.ai?campaign=grok-4-7-blog&utm_source=website&utm_medium=referral&utm_campaign=grok-4-7-blog), third-party coding harnesses, and model routers and cloud platforms.
The model is priced starting at $2 per million input tokens and $6 per million output tokens. We also serve a fast variant with twice the output speed at twice the price.
[Console
Create an API key](https://console.x.ai/team/default/api-keys?campaign=grok-4-7-blog&utm_source=website&utm_medium=referral&utm_campaign=grok-4-7-blog)[docs.x.ai
Read the docs](https://docs.x.ai)
### Try it in Grok Build for free
Get started today at [x.ai/build](https://x.ai/build).
`$ curl -fsSL https://x.ai/cli/install.sh | bash` | meetpateltech | 609 | 491 |
| 🟧 echo.blog ⭐ | Announces Grok 4.7 availability through Cursor, Grok Build, and the API; claims improved multi-hour task performance and frontier price-perf | SpaceXAI | — | — |
| 🟧 hn | Grok 4.7 Intelligence, Performance and Price Analysis | theanonymousone | 13 | 1 |
| 🟠 reddit | Elon Musk admits Grok isn’t as good as Anthropic’s Claude ClaudeAI | ross2000 | 769 | 117 |
2026-09-26T20:44:19Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-09-26T20:26:30Z
evidence attached: reddit.post.1wqzu96 — First-party admission by Musk that Grok trails Claude undercuts the case's frontier-comparable-performance claim.
2026-09-21T20:59:41Z
The purported independent analysis adds no usable benchmark results: its only supplied excerpt quotes implausible zero-dollar pricing, so it cannot corroborate the price-performance thesis. Attention nevertheless merits high heat given strong HN and Reddit spread and an additional analysis thread; the retrieved launch page also contradicts the cached grounding's claim that release and pricing remain undocumented.
2026-09-21T20:23:25Z
evidence attached: hn.story.49789558 — Independent Artificial Analysis benchmarking materially contextualizes Grok 4.7's claimed intelligence, performance, and price advantage.
2026-09-21T16:27:40Z
grounded: known/medium — The supported price/performance pitch adds no established position beyond Scott’s AI Unit Economics and Model-Plus-Harness Benchmark Unit: compare cost per succ
2026-09-21T16:24:08Z
case created — Both observations echo one available model release with concrete pricing and owner-reported benchmarks, warranting a single case rather than separate capability and release episodes.