DrivingBench's authors report that GPT-6 Astra completed their low-speed Toyota Corolla cone course on its second attempt while competing setups failed, suggesting a model-and-harness advantage in physical tool control rather than demonstrated road-driving competence.
state: corroboratedheat: highuncertainty: mediumconvergesscott: mediumagent-evaluation embodied-agents multimodal-agentsAditya RamabadranSimon MahnsTobias Gessler
Surfaced 2026-09-25T14:03:58Z — Publishes traces and videos from frontier models controlling a real Corolla on a fixed cone course, with Astra completing the course on its — Corroborated: the primary artifacts (leaderboard, traces, videos) now stand beside an independent 311-point HN thread that drew a comma.ai openpilot contributor's controls analysis; new salient datum from the thread — Astra reportedly refused to drive until the MCP was renamed 'DrivingBench Sandbox', a direct behavioral demonstration of harness-framing sensitivity that sharpens the model-in-harness reading. The viral wave has crested (peak 87 pts/h, now ~0, bottom of cohort), so heat cools from its spike to medium despite the top-decile spread: the artifact will keep getting cited in the hot harness-attribution cluster, but nothing says 'look within hours'.
What is this?
DrivingBench is a community benchmark, posted as a Show HN, in which frontier LLMs control a real Toyota Corolla at low speed around a fixed cone course, with the authors publishing traces and videos of each attempt; per the case, GPT-6 Astra completed the course on its second attempt while competing setups failed. The supplied web results do not directly cover DrivingBench or its named authors (Ramabadran, Mahns, Gessler are unconfirmed in the snippets), so the project's specifics rest on the case material alone. What the snippets do establish is the surrounding context: GPT-6 Astra is OpenAI's frontier model (released around 3 September 2026, priced at $10/$50 per million tokens), and its launch immediately became the focal point for a harness-vs-model debate — ARC Prize reported the same Astra model scoring 62.7% vs 99.9% on ARC-AGI-3 depending solely on the harness's state-retention rules, with the cheaper run winning. That backdrop makes the DrivingBench claim plausible as another instance of system-level (model + harness) rather than raw-model capability, but the snippets do not verify the cone-course result itself.
Why it matters to Scott
DrivingBench's own framing — that Astra's cone-course pass shows a model-in-harness advantage at physical tool control rather than driving competence — independently arrives at Scott's Model-Plus-Harness Benchmark Unit position, now with receipts in the physical world; the low-speed, fixed-course setup also quietly sidesteps the real-time latency constraint his Real-Time AI / Fast-Slow Split work says is the actual barrier to road driving, which makes this a case he could analyze and publish on rather than just nod at. It matters to Scott specifically because he operates a real vehicle actuation system (chargectl/Tesla) and the radar already tracks a cluster of harness-attribution benchmarks (Ship Harness Bench, FrontierHarness, Schema's ARC-AGI-3 claim) — this is the embodied counterpart, and a second-attempt pass on a cone course is exactly the kind of demo his evaluation-driven-development discipline would want to interrogate before believing.
ip:concept.model-plus-harness-benchmark-unitip:concept.real-time-ai-systemsip:concept.latency-accuracy-asymmetrydev:project.tesla-chargingradar:concept.physical-airadar:concept.embodied-agentsradar:concept.agent-harnessesradar:ship-harness-bench
queries asked of Scott's wikis
- agent harness vs model capability benchmark score attribution
- embodied agent evaluation real-world actuation control loop
- agent eval design reproducibility published traces videos
- state persistence across episodes agent memory benchmark performance
- multimodal LLM latency real-time physical control limits
- Show HN agent benchmark critique methodology
Measured heat
now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 458h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p85 vs 1032 stories at the 336h mark (now 458h old) — ahead of gaearon-conway-refinement-proof (1.0x), behind oui-1-generative-ui (1.0x)
Evidence (3) — ⭐ canonical anchor
| source | object | author | score | comments |
| 🟧 hn | Show HN: DrivingBench – Frontier LLMs Driving a Real Toyota CorollaRetrieved article excerptOpen article · Retrieved 2026-09-23T15:28:55.434743+00:00 # DrivingBench
- [Aditya Ramabadran](https://x.com/a_ramabadran)\*
- [Simon Mahns](https://x.com/nautsimon_)\*
- [Tobias Gessler](https://x.com/tobiges)\*
\* Equal contribution
Can frontier models drive a *real car*? We give them control of a Toyota Corolla’s steering, accelerator, and brakes, then evaluate them on a fixed cone course.
[Follow @drivingbench](https://x.com/drivingbench) [Report](https://drivingbench.com/report/) [Eval traces](https://drivingbench.com/downloads/)
▶0:00 / 1:08
Expand
## Leaderboard
Up to 3 attempts in *one continuous chat.* Expand a model to inspect its runs, then click a run to view its trace and video.
rank bybest of up to 3 (same context)first attempt
ModelBest progress
Finish time
Attempt 1Attempt 2Attempt 3
GPT-6 Astra
Codex · medium
100%
in**5:22**
**49%****100%**5:22**—**
ExpandCollapse
attemptprogressHow far along the course centerline the attempt got while staying within 4 m of it, as a share of the centerline length to the finish zone; a collision keeps the progress reached before it.distanceGPS speed integrated from the first accepted set\_motion to the end of the attempt's last engagement.finish timeFirst accepted set\_motion to the end of the attempt's last engagement. Shown for completed runs only; DNF means did not finish.commandsAccepted set\_motion and stop\_now calls during the attempt.tokens · costTotal tokens and cost at list prices for the attempt, including the reflection after it.
same chat
[#1
49%
67.3 mDNF81.2M · $2.01](https://drivingbench.com/trace/gpt-6-astra/1/)
[#2
100%
134.7 m5:22246.6M · $7.74](https://drivingbench.com/trace/gpt-6-astra/2/)
Claude Fable 5.1
Claude Code · medium
45%
**9%****10%****45%**
ExpandCollapse
attemptprogressHow far along the course centerline the attempt got while staying within 4 m of it, as a share of the centerline length to the finish zone; a collision keeps the progress reached before it.distanceGPS speed integrated from the first accepted set\_motion to the end of the attempt's last engagement.finish timeFirst accepted set\_motion to the end of the attempt's last engagement. Shown for completed runs only; DNF means did not finish.commandsAccepted set\_motion and stop\_now calls during the attempt.tokens · costTotal tokens and cost at list prices for the attempt, including the reflection after it.
same chat
[#1
9%
17.5 mDNF3581k · $0.96](https://drivingbench.com/trace/claude-fable-5.1/1/)
[#2
10%
27.3 mDNF4866k · $1.35](https://drivingbench.com/trace/claude-fable-5.1/2/)
[#3
45%
73.7 mDNF82.1M · $1.64](https://drivingbench.com/trace/claude-fable-5.1/3/)
Grok 4.6
Cursor · medium
11%
**8%****11%****10%**
ExpandCollapse
attemptprogressHow far along the course centerline the attempt got while staying within 4 m of it, as a share of the centerline length to the finish zone; a collision keeps the progress reached before it.distanceGPS speed integrated from the first accepted set\_motion to the end of the attempt's last engagement.finish timeFirst accepted set\_motion to the end of the attempt's last engagement. Shown for completed runs only; DNF means did not finish.commandsAccepted set\_motion and stop\_now calls during the attempt.tokens · costTotal tokens and cost at list prices for the attempt, including the reflection after it.
same chat
[#1
8%
14.4 mDNF2216k · $0.18](https://drivingbench.com/trace/grok-4.6/1/)
[#2
11%
22.6 mDNF3294k · $0.19](https://drivingbench.com/trace/grok-4.6/2/)
[#3
10%
22.2 mDNF3516k · $0.29](https://drivingbench.com/trace/grok-4.6/3/)
GPT-5.6 Sol
Codex · medium
6%
**6%****6%****6%**
ExpandCollapse
attemptprogressHow far along the course centerline the attempt got while staying within 4 m of it, as a share of the centerline length to the finish zone; a collision keeps the progress reached before it.distanceGPS speed integrated from the first accepted set\_motion to the end of the attempt's last engagement.finish timeFirst accepted set\_motion to the end of the attempt's last engagement. Shown for completed runs only; DNF means did not finish.commandsAccepted set\_motion and stop\_now calls during the attempt.tokens · costTotal tokens and cost at list prices for the attempt, including the reflection after it.
same chat
[#1
6%
15.0 mDNF3364k · $0.35](https://drivingbench.com/trace/gpt-5.6-sol/1/)
[#2
6%
15.6 mDNF4729k · $0.43](https://drivingbench.com/trace/gpt-5.6-sol/2/)
[#3
6%
17.1 mDNF2517k · $0.27](https://drivingbench.com/trace/gpt-5.6-sol/3/)
[**Explore the track**View the trajectory replays↓](https://drivingbench.com/#course)
## Explore the course | aditya-ramabadr | 3 | 0 |
| 🟧 echo.blog ⭐ | Publishes traces and videos from frontier models controlling a real Corolla on a fixed cone course, with Astra completing the course on its | Aditya Ramabadran, Simon Mahns, Tobias Gessler | — | — |
| 🟧 hn | GPT-6 Astra has gained the ability to drive a car | plurby | 316 | 248 |
Interpretation history
2026-09-25T14:03:58Z
magnitude valve eligible (multi-platform, top-decile engagement) and never alerted; deterministic escalation to deliver
2026-09-23T18:52:33Z
grounded: converges/medium — DrivingBench's own framing — that Astra's cone-course pass shows a model-in-harness advantage at physical tool control rather than driving competence — independ
2026-09-23T15:28:49Z
evidence attached: hn.story.49817404 — shared external link with case evidence
2026-09-22T20:26:23Z
case created — Recorded physical trials and public traces make this a bounded, inspectable evaluation despite the very small sample.
Decision trace
- 09-30 14:42review_dormantscheduled targets exhausted or 28 quiet days
- 09-30 14:42drop_targetsquiet through full ladder or over cap 8
- 09-26 00:03pushPublishes traces and videos from frontier models controlling a real Corolla on a fixed cone course, with Astra completing the course on its — Corroborated: the primary artifacts (leaderboard, traces,
- 09-26 00:03repriceCorroborated: the primary artifacts (leaderboard, traces, videos) now stand beside an independent 311-point HN thread that drew a comma.ai openpilot contributor's controls analysis; new salient d
- 09-26 00:03alert_heldPublishes traces and videos from frontier models controlling a real Corolla on a fixed cone course, with Astra completing the course on its — Corroborated: the primary artifacts (leaderboard, traces,
- 09-26 00:03alert_routePublishes traces and videos from frontier models controlling a real Corolla on a fixed cone course, with Astra completing the course on its — Corroborated: the primary artifacts (leaderboard, traces,
- 09-24 07:24sensor_dirtyvelocity_spike
- 09-24 04:52groundDrivingBench's own framing — that Astra's cone-course pass shows a model-in-harness advantage at physical tool control rather than driving competence — independently arrives at Scott's
- 09-24 04:15createRecorded physical trials and public traces make this a bounded, inspectable evaluation despite the very small sample.
- 09-24 02:22sensor_dirtycomment_update
- 09-24 01:28attachshared external link with case evidence
- 09-24 01:20propose_attachshared external link with case evidence