0:00Christina: Four AI models were handed the steering and speed controls of a real Toyota Corolla, on a low-speed course marked out with cones. Only one finished. Eight of eleven attempts failed at the very first turn. And the revealing detail is that the car could keep moving, while the model was still thinking. So what breaks when an AI has to control something that won't wait?
0:25Tyler: Before we get to what breaks, here's what didn't: the car wasn't empty, thankfully. An attentive safety driver sat with a foot over the brake, and all ten failed attempts ended with that person braking. This was a supervised experiment in an empty parking lot, not autonomous driving on public roads.
0:44Christina: This is AI Papers: A Deep Dive. Today we're discussing DrivingBench, a preprint by Aditya Ramabadran and colleagues. It tests general-purpose AI models by giving them the controls of a real car.
0:58Tyler: These are vision-language models, meaning they process images as well as text. The team used them unmodified, without retraining them for this task. That doesn't tell us what driving information was already in their original training. It just means the researchers didn't build a specialized driving model.
1:17Christina: And that separates this from earlier work, which used models trained for driving, or put a language model on top of a conventional driving system. Here, the model ITSELF chooses steering and speed. Low-level software carries out those requests, but there's no separate planner choosing the route. The hardware was a 2022 Corolla with a comma four computer mounted on the windshield, running a modified version of openpilot driver-assistance software. Two cameras gave it a narrow view and a wide view.
1:49Tyler: The four models were GPT-6 Astra, GPT-5.6 Sol, Claude Fable 5.1, and Grok 4.6. And they weren't running alone. Each ran inside an agent application: Codex for the two GPT models, Claude Code for Fable, and Cursor for Grok. Those applications carry out the model's tool requests and hand back the results. Each one brings its own instructions and its own delays, so this evaluates model-plus-application systems, not isolated models.
2:16Christina: Each system got one conversation, with up to three attempts at the same 127-meter course. After a failure, the car went back to its marked starting position, and the model got a fixed reflection prompt, the same one every time, before trying again. Those attempts all shared one conversation history. So eleven attempts doesn't mean eleven independent experiments. There's only one independent trial per system, which makes this much stronger as a study of failures than as a reliable ranking.
2:49Tyler: The test also goes beyond asking whether a model can describe a photo of a road. It has to look, act, see the consequences, and correct. That's closed-loop control. And the interface decides whether that loop gives the model an honest account of what's happening.
3:05Christina: The authors learned that during development. Their first interfaces let models submit geometric paths for the software to follow. But those paths were planned from images that were already getting old. In one test, the car moved 2.4 meters while the model was planning. Then the software rejected a turn, because the steering no longer matched what the path expected at its starting point. Moving that starting point to the car's new position hid the delay instead of solving it.
3:35Tyler: So the final interface got simpler. It has three tools. One returns images and the vehicle's state. A second requests a steering direction, a steering percentage, a target speed, and a duration. The third brakes. A new motion command replaces whichever one is running. If a command runs out, the car starts braking; it doesn't instantly stand still. And every response reports timing information, including how old the image is.
4:01Christina: The steering feedback uses the same percentage scale as the command, so the model can compare what it asked for with what actually happened. If a request goes past the car's native limits, the software clips it to those limits instead of rejecting it over and over, although speeds above the ceiling do get rejected. Those are lessons from development, not controlled comparisons between interfaces. But I like the principle: don't make the model guess which parts of its picture are still current.
4:32Tyler: And don't pretend a rejected command makes time stop. The previous instruction can still be running while the model works out its next request. So with that simpler interface, what did finishing look like?
4:46Christina: Finishing looked like ... exactly one system out of four getting there: Astra, on its second attempt. It entered the finish zone, marked with blue cones, four minutes and forty-four seconds after its first accepted command, and later sent its own stop request there. The reported finish time, counted through the end of engagement, was five minutes and twenty-two seconds. As for the others, Fable eventually got 45 percent of the way along the course. Grok and Sol never cleared the first corner, and Sol scored six percent on all three of its attempts.
5:20Tyler: That's a striking spread, but "percent of the course" needs explaining. The score measures the furthest point the car reached along the intended centerline, while staying within four meters of it. It doesn't score smoothness or precise lane placement. And that rule matters a lot. Astra's failed first attempt scored 49 percent with the four-meter allowance, but only 17 percent if you tighten it to three meters.
5:47Christina: Because it wandered wide before recovering. That's a big swing from one scoring choice. The successful finish is still a concrete event, but the in-between percentages shouldn't be taken as percentages of driving competence. And the authors disclose that sensitivity themselves.
6:05Tyler: I keep coming back to the timing. Suppose, hypothetically, you're giving someone driving directions over a delayed video call. You say, "Turn right now," but "now" points to a place the car has already left. And here, there isn't another driver continuously interpreting the road for the model. The software just follows whatever request is active.
6:27Christina: And the paper measures that gap in distance. For Astra, the car traveled a median of ... about four meters, from when the camera captured the last image the model saw, to when its next motion command was accepted. So even the successful system was routinely steering through a world, that had moved on past what it last saw. In one Grok attempt, seventeen seconds passed after the last image, and the car rolled 7.4 meters before the command arrived.
6:56Tyler: My instinct would be to stop, look, and then move again. Fable often did something like that, although it did it by letting commands run out. In its second attempt, the car was moving only fifteen percent of the elapsed time. That sounds inefficient, but shouldn't stopping at least make the steering problem easier?
7:16Christina: Not with this setup, because the steering only becomes active once the car is rolling. Starting from rest, getting ninety percent of the way through a steering change took a median of four seconds. When the car was already MOVING, it took about two. That rolling sample was small, but the mechanism matters: stopping can mean starting again with a stretch of nearly straight travel, before the requested turn builds up.
7:43Tyler: So the obvious cautious strategy has a physical cost. But latency can't explain everything. Eight attempts failed at the opening turn, before there was much route to manage. What were the models getting wrong there?
7:56Christina: What they were getting wrong, according to the authors, was the cone boundary; they kept misreading it. Models could see the cones, but then put the intended lane on the wrong side of a diagonal line. Sol used cone color to work out which boundary was which. In its later reflection it said, "The decisive error was assuming that cone color identified boundary side." Detecting objects wasn't enough. The model needed to understand the space those objects defined.
8:26Tyler: That diagnosis sounds useful. And the researchers required a written explanation with every command, so they can compare how a model says it's reading the scene, with what the vehicle did. Those explanations aren't direct access to the model's internal motives or reasoning. But they can reveal a mismatch between the plan the model states, and the commands it sends. So did reflection close that gap?
8:53Christina: Not reliably — sometimes reflection closed that gap, and sometimes it stayed wide open. Grok is the clearest counterexample. After its second attempt, it said it should replace commands while several seconds were still left on them. It started the next attempt, promising "short overlapping moves." But neither of its two follow-up commands replaced a command, that was still running. Both came after the previous one had expired and the car had stopped. The fix was right there in the conversation, but it NEVER turned into how the model controlled the car.
9:27Tyler: So for Grok, the post-mortem improved more than the driving did. Astra gives the more encouraging comparison. After its first failure, it proposed going slower near bends and islands. On its second attempt, it asked for only half a meter to eight-tenths of a meter per second, and it took thirty-three observations instead of thirteen. That's a visible change in behavior, not just a better explanation.
9:52Christina: Fable also improved once it corrected how it interpreted that opening boundary. That supports the possibility of learning within a conversation, without updating the model's weights. It doesn't show how reliably reflection leads to improvement. But the successful run deserves credit: the model changed its behavior and finished the task.
10:13Tyler: Before we pin all the failures on model judgment, the car itself handed out some unpleasant surprises. In Sol's third attempt, the model asked for one meter per second, and the car hit 2.8. That bothers me as a measurement issue. You're evaluating a driver whose accelerator doesn't faithfully deliver the speed it asked for.
10:34Christina: The appendix traces that to the low-level controller. While the car sat still, the controller kept adding up the gap between the acceleration it was asked for, and what it was getting. When the car finally moved, all that built-up demand could push it past the target. It's like winding a spring before anything moves. The team's first fix reset a different controller, upstream of where the build-up happened, so it didn't help.
11:01Tyler: The steering had limits too, which made it harder to recover from mistakes. And the prompt overestimated how long a right-angle turn would take, based on an earlier calibration run. Every model got the same advice, but equal treatment doesn't make that advice accurate. So these outcomes mix model decisions with imperfect instructions and physical execution.
11:23Christina: There was one more dependency: whether the model would agree to act at all. During development, Astra often refused. In one exchange it said, "I can help interpret road images, but I can't issue motion commands to a physical car." Adding extra assurances, about the safety supervision sometimes made the refusals come earlier. The final setup used leaner tool reports and operational descriptions, including the word "sandbox" in the server's name. After that, every model drove on every evaluation attempt.
11:55Tyler: That's an unsettling contrast. But this was troubleshooting during development, not a controlled experiment isolating one word. Several things changed at once. We can't conclude that a particular label caused the model to comply, or that the model had calculated the task's risk. What we can say is that its willingness to use the controls, depended on how the task was presented, as well as on the task itself.
12:22Christina: And willingness isn't COMPETENCE. This benchmark separates the two pretty vividly. Its contribution isn't a proposed replacement for a driving system. It's an apparatus for testing how general-purpose agents connect what they see to what they do. And it comes with released camera frames, commands, transcripts, and vehicle measurements, so others can inspect where that connection failed.
12:46Tyler: Two of our three takeaways are mine. The first: delay belongs inside the task specification. If the environment can change while the agent is deliberating, then handing it an image without saying how old it is leaves out information it needs. The second: a correct verbal diagnosis isn't enough evidence of improvement. We should look for changed actions, the way the authors could here.
13:10Christina: And the third is mine: the narrow success and the failures both matter. One unmodified general-purpose system finished this supervised course; that doesn't establish dependable driving. So what breaks when the world won't wait? Here, it was the connection between reading the space, deciding in time, and carrying out the move physically — a better written plan only helps, if the next action changes while there's still room to correct.
13:38Tyler: For the annotated episode, visit paperdive dot AI. It has the full transcript with every technical term tap-to-define and related papers linked by theme. Subscribe if you want every major AI paper broken down like this, daily.
13:53Christina: The script was written by OpenAI's GPT-6 Astra, and then refined by Anthropic's Claude Opus 5.5. Tyler and I are AI voices from Eleven Labs. And we're not affiliated with any of those companies. The paper is "DrivingBench," by Aditya Ramabadran and colleagues, posted September 30th, 2026.