Stanford's HomeBody lets GPT-6 Astra run a humanoid robot with no trained control layer
Most humanoid robots reason with one model and move with a second one trained specifically to translate that reasoning into motion. A Stanford and Caltech team deleted the second model entirely and had a Unitree G1 robot tidy an unfamiliar kitchen anyway, guided step by step by OpenAI's GPT-6 Astra.
What
Researchers from Stanford's Movement Lab and Caltech published HomeBody around September 26 to 27, 2026, describing a system that skips the learned vision-language-action policy that normally sits between a robot's high-level reasoning model and its low-level motor controller, per the project page. Instead, GPT-6 Astra calls directly into a library of five swappable skills (navigate, pick, place, open drawer, and pick from a drawer) selecting a target and a hand for each action and revising its plan when a skill reports failure. Before acting, the G1 explores the room using an iPhone camera, a D435i stereo camera, and LiDAR, then uses that data to build a digital twin inside Nvidia's Isaac Sim so Astra can reason about objects outside its current view, according to the project page. In demonstrations, the robot gathered coffee bags, discarded spoiled milk and juice cartons, and retrieved medicine from a closed drawer it had to locate from memory. The perception and motion-planning stack runs locally on a single laptop with an RTX 4090 GPU while Astra runs remotely over the network, the researchers wrote. The code is public on GitHub.
A trained middle layer may not be mandatory for humanoid autonomy
If a frontier vision-language model can plan and self-correct against a fixed skill library, robotics teams may not need to collect the environment-specific training data that a learned action policy normally requires. That matters for anyone evaluating humanoid platforms: the bottleneck shifts from training a new policy per robot or per task toward building and maintaining a shared skill interface that any capable VLM can call. It also means the underlying reasoning model becomes swappable the way the researchers designed it, so a robot's capabilities could improve simply by upgrading the model behind it rather than retraining its controller.
The approach has real limits today
Stanford's own writeup names the friction points: building the Real2Sim digital twin adds setup time and API cost, Astra's reasoning latency introduces pauses between each skill, and extended operation risks overheating the robot's finger servos, per the project page. The Decoder also noted a separate benchmark had flagged safety issues when Astra and Anthropic's Claude Fable controlled robot arms, per its report, a reminder that giving a language model direct control over hardware raises questions this architecture does not resolve on its own.
What to watch next
OpenAI has said it wants to get back into robotics infrastructure with an eventual goal of a personal robot for everyday tasks, according to The Decoder, and HomeBody's open-source skill library gives outside teams a concrete way to test whether the VLA-free approach generalizes past kitchen chores to other household or warehouse tasks.
Sources
- HomeBody: A Humanoid That Explores, Remembers, and Acts on Its Own: Stanford Movement Lab, primary
- Researchers plug GPT-6 Astra directly into a robot and let it clean up an unfamiliar kitchen: The Decoder, secondary
