GPT-6 Astra just drove a car around a parking lot, covering 134.7 meters of cone-marked course in 5 minutes and 22 seconds while three rival models largely failed before the first bend. Here is what that actually means.
DrivingBench, created by researchers Aditya Ramabadran, Simon Mahns, and Tobias Gessler, is one of the first structured attempts to test whether general-purpose agents can handle real physical control rather than just answer questions fluently.
What the Experiment Actually Tested
Four frontier AI models were wired into a real Toyota, handed the controls, and told to drive a simple cone-marked course.
The researchers connected GPT-6 Astra, GPT-5.6 Sol, Claude Fable 5.1, and Grok 4.6 to a 2022 Toyota Corolla. The car ran Comma Four hardware and openpilot control software.
Each model received live camera imagery and vehicle telemetry, then issued steering, acceleration, and braking commands over a wireless link to remote data centers. At each step it narrated its observations and intentions.
That wireless link matters. Because every command traveled to a data center and back, the car regularly paused mid-maneuver waiting for the next instruction, which is not how purpose-built autonomous systems are designed to operate.

Who Finished, Who Didn’t
Of eleven total runs across four models, only one completed the course.
Astra was the only model to finish, doing so on its second attempt. Its path hugged a long right-hand bend too closely and nearly left the course near the finish line, but it reached the endpoint.
Claude Fable 5.1 reached roughly 45% of the course at its best. Grok 4.6 reached approximately 11%, with its first attempt ending after just two commands: it read a gap between cones as a gate and drove straight toward it. GPT-5.6 Sol managed approximately 6%.
The cost figures make the result more concrete. Astra’s successful run consumed approximately 6.6 million tokens and cost $7.74. That is nearly four times the cost of its previous attempt, which covered about half the distance.
Completing a physical task through a language model is less like asking for directions and more like texting a friend turn-by-turn instructions from the backseat, with a noticeable delay between every message.
What the Failures Reveal
Fluent reasoning and reliable physical control turned out to be very different capabilities.
The errors went beyond misunderstood instructions. Models struggled with spatial reasoning, specifically translating a desired turn into a steering angle that actually matched it. One GPT model invented a cone-color rule that the instructions explicitly contradicted.
Closed-loop control exposed the deepest problem: the repeated cycle of observing the vehicle’s state, acting, and correcting based on the result revealed how poorly fluent language generation maps onto consistent physical motion.
You can watch a model narrate exactly what it sees while the car drifts the wrong direction.
DrivingBench did not test against commercial autonomous-driving systems, which combine specialized perception, localization, prediction, planning, and layered safety architecture. This experiment connected general-purpose agents to a real vehicle through an experimental harness, and that distinction matters for interpreting what the results show.
The Takeaway
One parking-lot lap is not a road test, but the methodology behind it is worth tracking.
One successful low-speed run does not establish driving competence, repeatability, or readiness for public roads. What it does suggest is that future AI evaluations may increasingly ask whether agents can handle physical tasks, not just text. The gap between explaining what you see and reliably controlling what you touch remains wide.




























