TL;DR
Get garage and car supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
DrivingBench researchers tested general-purpose AI models using Comma vision hardware and OpenPilot to steer a Toyota Corolla around a cone-marked parking-lot course. Eight of 11 runs across four agents failed to complete more than 11% of the route; GPT-6 Astra completed it on its second attempt, while Grok drove off course on its first.
Researchers behind DrivingBench tested whether general-purpose AI models could steer a Toyota Corolla around a cone-marked parking-lot course, and most attempts failed near the start. Across 11 runs involving four agents, eight completed no more than 11% of the route; GPT-6 Astra was the only agent reported to finish, doing so on its second attempt.
The researchers, identified in The Drive’s report as Aditya Ramabadran, Simon Mahns and Tobias Gessler, used Comma vision hardware and OpenPilot to give the agents control of the car. The course was a short point-to-point route laid out with small cones, including several curves. Each agent received the same instructions and was asked to describe what it saw and did after each step.
The test involved 11 runs across four agents, with up to three attempts for each. Because commands were sent remotely over a wireless hotspot, the car regularly stopped between actions. The trials were conducted within continuous chat sessions, and DrivingBench published videos showing the vehicle’s movements alongside the agents’ messages.
The report describes several different failure modes. Grok treated a gap between boundary cones as a gate and drove off course on its first attempt, which ended after two commands. A GPT model misread the cone layout and invented a color rule despite the written instructions warning that the cones were not all the same color. Other agents did not turn far enough to follow the bends. GPT-6 Astra eventually completed the route on its second attempt, but its path was uneven and came close to leaving the course near the finish.
Why General AI Struggled at Driving
The test illustrates a gap between handling language-based tasks and making reliable decisions in a physical environment. A driving agent must interpret camera input, understand the route, judge how far to turn and respond to the vehicle’s movement. In these trials, models made errors in reading basic course features and translating their plans into steering actions.
The findings are relevant to claims that general-purpose AI could readily take on real-world tasks, but they are not a test of every autonomous-driving system. The experiment used a small course and a particular hardware and software setup, and the report does not establish how these agents would perform in other conditions. It does show that the tested models were not consistently able to complete even this limited route.
As an affiliate, we earn on qualifying purchases.
How DrivingBench Set Up the Test
The report distinguishes the tested models from purpose-built self-driving systems. DrivingBench asked general-purpose agents, including Claude, GPT and Grok models, to operate through Comma vision hardware and OpenPilot. The point was to examine how these agents handled driving instructions and visual information, rather than to assess a dedicated commercial driver-assistance or autonomous-driving product.
The trials took place on a short, cone-defined parking-lot course, not on public roads. The researchers gave agents matching instructions, allowed multiple attempts and recorded the chats and vehicle paths. The Drive’s account says the successful GPT-6 Astra run cost $7.74, nearly four times the charge for its previous attempt, which had covered about half as much of the course. That cost refers to the reported run, not to a general estimate for AI driving.
“Of the 11 runs conducted across four different agents, eight failed to complete more than 11% of the course.”
— The Drive, summarizing the results
As an affiliate, we earn on qualifying purchases.
Limits of the Parking-Lot Results
The source account does not provide enough detail to establish how the results would generalize beyond this single course and test setup. It does not report trials on public roads, performance in different weather or traffic, or comparisons with trained human drivers and purpose-built autonomous systems. The total of 11 runs is also limited, so the results should not be treated as a broad ranking of all models.
The report does not specify every model version, the full technical configuration, or the complete scoring method in the material provided. It is also unclear whether the agents had any training or adaptation specific to the course. GPT-6 Astra’s completion establishes that it finished this particular test route once; it does not show that it can drive safely or reliably in general.
As an affiliate, we earn on qualifying purchases.
Further Testing Needed
The immediate next step for readers seeking a fuller picture is to examine the DrivingBench trial videos and route records, which the report says show the chats and vehicle paths. Further runs would be needed to test repeatability, use more varied courses, document model and hardware configurations, and compare results under consistent scoring.
No follow-up schedule or broader testing plan is specified in the source material. Until such evidence is available, the findings are best read as a limited demonstration of how four general-purpose agents handled one controlled course—not as evidence that these models are ready to take over driving.
Toyota Corolla parking lot navigation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Did ChatGPT successfully drive the course?
GPT-6 Astra, a GPT model identified in the report, completed the cone-marked route on its second attempt. The report says its path was uneven and came close to leaving the course near the finish.
What happened when Grok tried the course?
According to The Drive’s account, Grok mistook a gap between boundary cones for a gate and drove off course on its first attempt. That run ended after two commands.
How many attempts failed early?
Of 11 runs across four agents, eight failed to complete more than 11% of the route, according to the report. The figure describes those test runs, not a general failure rate for AI systems.
Was this a test of autonomous cars on public roads?
No. The test used a short parking-lot course marked with cones, a Toyota Corolla, Comma vision hardware and OpenPilot. The report does not describe public-road testing.
Does the experiment show that general AI is ready to drive?
No. One model completed this particular route, while most runs failed early. The limited trial does not establish safe, repeatable performance in other settings or conditions.
Source: rss
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
