2026-10-11 16:38 UTC

DrivingBench's authors report that GPT-6 Astra completed their low-speed Toyota Corolla cone course on its second attempt while competing setups failed, suggesting a model-and-harness advantage in physical tool control rather than demonstrated road-driving competence.

state: corroboratedheat: highuncertainty: mediumconvergesscott: mediumagent-evaluation embodied-agents multimodal-agentsAditya RamabadranSimon MahnsTobias Gessler
Surfaced 2026-09-25T14:03:58Z — Publishes traces and videos from frontier models controlling a real Corolla on a fixed cone course, with Astra completing the course on its — Corroborated: the primary artifacts (leaderboard, traces, videos) now stand beside an independent 311-point HN thread that drew a comma.ai openpilot contributor's controls analysis; new salient datum from the thread — Astra reportedly refused to drive until the MCP was renamed 'DrivingBench Sandbox', a direct behavioral demonstration of harness-framing sensitivity that sharpens the model-in-harness reading. The viral wave has crested (peak 87 pts/h, now ~0, bottom of cohort), so heat cools from its spike to medium despite the top-decile spread: the artifact will keep getting cited in the hot harness-attribution cluster, but nothing says 'look within hours'.

What is this?

DrivingBench is a community benchmark, posted as a Show HN, in which frontier LLMs control a real Toyota Corolla at low speed around a fixed cone course, with the authors publishing traces and videos of each attempt; per the case, GPT-6 Astra completed the course on its second attempt while competing setups failed. The supplied web results do not directly cover DrivingBench or its named authors (Ramabadran, Mahns, Gessler are unconfirmed in the snippets), so the project's specifics rest on the case material alone. What the snippets do establish is the surrounding context: GPT-6 Astra is OpenAI's frontier model (released around 3 September 2026, priced at $10/$50 per million tokens), and its launch immediately became the focal point for a harness-vs-model debate — ARC Prize reported the same Astra model scoring 62.7% vs 99.9% on ARC-AGI-3 depending solely on the harness's state-retention rules, with the cheaper run winning. That backdrop makes the DrivingBench claim plausible as another instance of system-level (model + harness) rather than raw-model capability, but the snippets do not verify the cone-course result itself.

Why it matters to Scott

DrivingBench's own framing — that Astra's cone-course pass shows a model-in-harness advantage at physical tool control rather than driving competence — independently arrives at Scott's Model-Plus-Harness Benchmark Unit position, now with receipts in the physical world; the low-speed, fixed-course setup also quietly sidesteps the real-time latency constraint his Real-Time AI / Fast-Slow Split work says is the actual barrier to road driving, which makes this a case he could analyze and publish on rather than just nod at. It matters to Scott specifically because he operates a real vehicle actuation system (chargectl/Tesla) and the radar already tracks a cluster of harness-attribution benchmarks (Ship Harness Bench, FrontierHarness, Schema's ARC-AGI-3 claim) — this is the embodied counterpart, and a second-attempt pass on a cone course is exactly the kind of demo his evaluation-driven-development discipline would want to interrogate before believing.
ip:concept.model-plus-harness-benchmark-unitip:concept.real-time-ai-systemsip:concept.latency-accuracy-asymmetrydev:project.tesla-chargingradar:concept.physical-airadar:concept.embodied-agentsradar:concept.agent-harnessesradar:ship-harness-bench
queries asked of Scott's wikis
  • agent harness vs model capability benchmark score attribution
  • embodied agent evaluation real-world actuation control loop
  • agent eval design reproducibility published traces videos
  • state persistence across episodes agent memory benchmark performance
  • multimodal LLM latency real-time physical control limits
  • Show HN agent benchmark critique methodology

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 458h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-22 20:26 (minted)⭐ origin echo-reconstructedPublishes traces and videos from frontier models controlling a real Corolla on a fixed cone course, with Astra completing the course on its
Aditya Ramabadran, Simon Mahns, Tobias Gessler on blog (echo) · attributed from hn.story.49801925 · published time unknown
—
09-22 14:29first on hacker news · published · lag ?Show HN: DrivingBench – Frontier LLMs Driving a Real Toyota Corolla
aditya-ramabadr
—
09-22 14:29amplified on hacker newshn.story.49801925
aditya-ramabadr
peak 3 · 0 comments · 1% of case engagement
09-23 15:14amplified on hacker news 👑hn.story.49817404
plurby
peak 316 · 248 comments · 99% of case engagement
09-22 15:20our radar first saw it · lag ?discovery anchor: hn.story.49801925—
09-25 14:03reached heat=high · lag ? · via queue+ledger——
pace: p85 vs 1032 stories at the 336h mark (now 458h old) — ahead of gaearon-conway-refinement-proof (1.0x), behind oui-1-generative-ui (1.0x)

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: DrivingBench – Frontier LLMs Driving a Real Toyota Corolla
Retrieved article excerpt

Open article · Retrieved 2026-09-23T15:28:55.434743+00:00

# DrivingBench

- [Aditya Ramabadran](https://x.com/a_ramabadran)\*
- [Simon Mahns](https://x.com/nautsimon_)\*
- [Tobias Gessler](https://x.com/tobiges)\*

\* Equal contribution

Can frontier models drive a *real car*? We give them control of a Toyota Corolla’s steering, accelerator, and brakes, then evaluate them on a fixed cone course.

[Follow @drivingbench](https://x.com/drivingbench) [Report](https://drivingbench.com/report/) [Eval traces](https://drivingbench.com/downloads/)

▶0:00 / 1:08

Expand

## Leaderboard

Up to 3 attempts in *one continuous chat.* Expand a model to inspect its runs, then click a run to view its trace and video.

rank bybest of up to 3 (same context)first attempt

ModelBest progress  
Finish time

Attempt 1Attempt 2Attempt 3

GPT-6 Astra

Codex · medium

100%

in**5:22**

**49%****100%**5:22**—**

ExpandCollapse

attemptprogressHow far along the course centerline the attempt got while staying within 4 m of it, as a share of the centerline length to the finish zone; a collision keeps the progress reached before it.distanceGPS speed integrated from the first accepted set\_motion to the end of the attempt's last engagement.finish timeFirst accepted set\_motion to the end of the attempt's last engagement. Shown for completed runs only; DNF means did not finish.commandsAccepted set\_motion and stop\_now calls during the attempt.tokens · costTotal tokens and cost at list prices for the attempt, including the reflection after it.

same chat

[#1

49%

67.3 mDNF81.2M · $2.01](https://drivingbench.com/trace/gpt-6-astra/1/)

[#2

100%

134.7 m5:22246.6M · $7.74](https://drivingbench.com/trace/gpt-6-astra/2/)

Claude Fable 5.1

Claude Code · medium

45%

**9%****10%****45%**

ExpandCollapse

attemptprogressHow far along the course centerline the attempt got while staying within 4 m of it, as a share of the centerline length to the finish zone; a collision keeps the progress reached before it.distanceGPS speed integrated from the first accepted set\_motion to the end of the attempt's last engagement.finish timeFirst accepted set\_motion to the end of the attempt's last engagement. Shown for completed runs only; DNF means did not finish.commandsAccepted set\_motion and stop\_now calls during the attempt.tokens · costTotal tokens and cost at list prices for the attempt, including the reflection after it.

same chat

[#1

9%

17.5 mDNF3581k · $0.96](https://drivingbench.com/trace/claude-fable-5.1/1/)

[#2

10%

27.3 mDNF4866k · $1.35](https://drivingbench.com/trace/claude-fable-5.1/2/)

[#3

45%

73.7 mDNF82.1M · $1.64](https://drivingbench.com/trace/claude-fable-5.1/3/)

Grok 4.6

Cursor · medium

11%

**8%****11%****10%**

ExpandCollapse

attemptprogressHow far along the course centerline the attempt got while staying within 4 m of it, as a share of the centerline length to the finish zone; a collision keeps the progress reached before it.distanceGPS speed integrated from the first accepted set\_motion to the end of the attempt's last engagement.finish timeFirst accepted set\_motion to the end of the attempt's last engagement. Shown for completed runs only; DNF means did not finish.commandsAccepted set\_motion and stop\_now calls during the attempt.tokens · costTotal tokens and cost at list prices for the attempt, including the reflection after it.

same chat

[#1

8%

14.4 mDNF2216k · $0.18](https://drivingbench.com/trace/grok-4.6/1/)

[#2

11%

22.6 mDNF3294k · $0.19](https://drivingbench.com/trace/grok-4.6/2/)

[#3

10%

22.2 mDNF3516k · $0.29](https://drivingbench.com/trace/grok-4.6/3/)

GPT-5.6 Sol

Codex · medium

6%

**6%****6%****6%**

ExpandCollapse

attemptprogressHow far along the course centerline the attempt got while staying within 4 m of it, as a share of the centerline length to the finish zone; a collision keeps the progress reached before it.distanceGPS speed integrated from the first accepted set\_motion to the end of the attempt's last engagement.finish timeFirst accepted set\_motion to the end of the attempt's last engagement. Shown for completed runs only; DNF means did not finish.commandsAccepted set\_motion and stop\_now calls during the attempt.tokens · costTotal tokens and cost at list prices for the attempt, including the reflection after it.

same chat

[#1

6%

15.0 mDNF3364k · $0.35](https://drivingbench.com/trace/gpt-5.6-sol/1/)

[#2

6%

15.6 mDNF4729k · $0.43](https://drivingbench.com/trace/gpt-5.6-sol/2/)

[#3

6%

17.1 mDNF2517k · $0.27](https://drivingbench.com/trace/gpt-5.6-sol/3/)

[**Explore the track**View the trajectory replays↓](https://drivingbench.com/#course)

## Explore the course
aditya-ramabadr30
🟧 echo.blog ⭐Publishes traces and videos from frontier models controlling a real Corolla on a fixed cone course, with Astra completing the course on its Aditya Ramabadran, Simon Mahns, Tobias Gessler——
🟧 hnGPT-6 Astra has gained the ability to drive a carplurby316248

Interpretation history

Decision trace