Back
GPT-6 Astra completes the DrivingBench cone course in 5:22, the only model to do so
SiTech AI Team3 წთ. საკითხავი

GPT-6 Astra completes the DrivingBench cone course in 5:22, the only model to do so

DrivingBench handed four frontier models control of a 2022 Toyota Corolla on a fixed cone course in an empty parking lot. GPT-6 Astra was the only one to finish the route, reaching the blue zone in 5 minutes 22 seconds.

DrivingBench, a benchmark built by Aditya Ramabadran, Simon Mahns and Tobias Gessler, asks whether frontier language models can drive a real car. The team laid out a cone course in an empty parking lot and gave four models control of a 2022 Toyota Corolla. Only one of them completed the route.

How the car is driven

A comma four device links the car to the benchmark. Models stay in their normal chat harnesses — Codex, Claude Code and Cursor — and call three MCP tools: observe(), which returns camera frames plus speed and steering; set_motion(direction, steering_percent, speed_mps, duration_s, reason); and stop_now(reason). A new command replaces the previous one instead of queuing, and the car keeps moving while the model thinks.

A human operator sits in the driver's seat with his foot over the brake, and the software caps speed at 0.5–3.5 m/s (1–8 mph), with an emergency stop that cancels motion above 6 m/s (13 mph).

DrivingBench's cone course in an empty parking lot

The leaderboard

Each model had up to three attempts in one continuous chat, with a reflection prompt between them. Progress counts how far along the course centreline the car travelled while staying within four metres of it; a collision keeps only the progress made before the impact.

GPT-6 Astra (Codex, medium) finished on its second attempt: 100% of the course in 5 minutes 22 seconds, across 24 commands. Its first attempt ended at 49%. Claude Fable 5.1 (Claude Code) peaked at 45% on its third attempt, Grok 4.6 (Cursor) at 11%, and GPT-5.6 Sol (Codex) at 6%. Every other attempt stopped before the first corner.

What the runs show

Most failures were perceptual rather than mechanical: models misread which side of the first diagonal cone line the lane was on. Astra observed roughly every five to six seconds; on the successful run it never exceeded 0.8 m/s and asked for full steering on 20 of its 24 commands.

The report also describes real in-context learning: Astra and Fable 5.1 both changed their steering and speed habits after reflecting on their mistakes.

What it does and does not show

This is a fixed cone course in an empty lot, not road autonomy. The authors list clear limits: each model was evaluated once, and the three attempts share one context, so they are not independent. They also note that some models refused to drive for safety reasons until the MCP server was renamed. DrivingBench is not affiliated with comma.ai, openpilot or Toyota. The team's conclusion is measured — that models can now drive a real car at low speed, and that the result calls for more work on safety and evaluation.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.