Text-Trained AI Model Steers a Car to an In-N-Out Pickup Window
Three Axiom engineers let OpenAI's GPT-6 Astra drive a Toyota Corolla up to a fast-food pickup window. The experiment suggests language models are gaining a basic grasp of the physical world.

Three AI engineers from the startup Axiom — Aditya Ramabadran, Simon Mahns and Tobias Gessler — recently tried an unusual way of getting lunch. Near an In-N-Out restaurant in the Bay Area, they sat in a 2024 Toyota Corolla, opened a laptop and asked OpenAI's GPT-6 Astra to drive.
The interface was linked to a server, which in turn connected to several cameras mounted on the windscreen and to the car's power steering. A safety driver kept a foot above the brake. The model, which normally produces text, code and images, slowly but successfully reached the pickup window. According to the engineers, it did so on the fly, with no prior training for the task.
Physical understanding as the next frontier
Today's most capable models answer complex questions well, but their skills are largely confined to computers and the internet. Some researchers have left big companies to found startups focused on physical reasoning. One is Elorian AI, whose CEO Andrew Dai previously worked at Google DeepMind. He argues that better visual reasoning will open new applications, such as home robots.
Elorian and Scale AI recently developed a benchmark called Humanity's Sixth Sense, which measures models' ability to understand physical scenes.
How the idea came about
The engineers ran the project outside their work at Axiom. They had noticed that models can build complex 3D simulations and wondered whether that would carry over to navigating the real world. At first the models refused to control the vehicle, but careful prompting coaxed them into doing so.
The engineers suspect the ability emerged from scaling up multimodal training. In their DrivingBench test, run on a parking-lot course, only Astra finished, and very slowly. Claude Fable 5.1 covered 45 percent of the course and Grok just 11 percent. Ramabadran said the models appeared to adapt to the controls and improve by learning from their mistakes.


