Back to list
Sep 28, 2026
941
0
0
ResearchNEW

Stanford HomeBody: A Frontier VLM Skips the VLA Layer

HomeBody lets GPT-6 Astra directly call a Unitree G1's skill library, skipping the usual learned VLA layer, to tidy an unfamiliar kitchen.

#HomeBody#GPT-6 Astra#Humanoid Robotics#VLM#Robotics Research
Stanford HomeBody: A Frontier VLM Skips the VLA Layer
AI Summary

HomeBody lets GPT-6 Astra directly call a Unitree G1's skill library, skipping the usual learned VLA layer, to tidy an unfamiliar kitchen.

Introduction

In September 2026, a Stanford and Caltech research team — Gio Huh, Cayden Gu, Takara E. Truong, C. Karen Liu, and Guy Tevet — published HomeBody, a humanoid robotics system built around a Unitree G1 and OpenAI's GPT-6 Astra. The project page frames a standard question in humanoid autonomy: most stacks chain a System 2 vision-language model (VLM) for reasoning, a learned System 1 vision-language-action (VLA) model that turns that reasoning into commands, and a System 0 whole-body controller that executes motion. HomeBody asks whether that middle, learned VLA layer is still needed once a frontier VLM is capable enough, and answers by removing it: GPT-6 Astra calls directly into a composable library of hand-built motor skills, with persistent spatial memory standing in for the learned layer's job of grounding decisions in the room. The project's paper link is marked "coming soon" on the project page, and the GitHub repository at Stanford-TML/homebody contains only a README and documentation figures — the README states "Code coming soon," so no runnable implementation is public.

Feature Overview

HomeBody's deployment follows two setup steps before any task runs. In the Explore phase, the humanoid collects 0.5x-speed iPhone video, Intel RealSense D435i camera frames, LiDAR scans processed through SLAM, joint poses, and a sequence of waypoints that Astra itself selects while exploring the space. In the Real2Sim phase, Astra acts as its own reconstruction agent, using that captured data to build a digital twin of the kitchen inside NVIDIA Isaac Sim. Stored keyframes from this twin let the system locate objects that are no longer in the robot's field of view, which is what lets it later retrieve something it saw earlier but can no longer see.

The system's "composable skill library" is a set of five real-robot skills documented on the project page: pick, place, navigate, open drawer, and pick-from-drawer. Each skill shares a common interface for targets and execution results, so the VLM selects a skill and a target — for picking, a normalized image point and a hand choice — through a structured tool call, without needing to know how the skill is implemented underneath. Segmentation and a stereo depth model (Fast-FoundationStereo) turn that point into a 3D grasp pose; a spline-based arm planner then solves inverse kinematics and checks for collisions. Locomotion runs through a pretrained whole-body controller (AMO) that coordinates the lower body with whatever the arms are doing, at 250 Hz for arm and hand commands and 50 Hz for the AMO policy itself. If a skill fails locally — for example, a grasp closes on nothing — it retries or adjusts before reporting the failure back up to the VLM, which can then choose a different target or plan.

Two demonstrations appear on the project page, both performed in a previously unseen kitchen without any environment-specific training data or additional policy learning. In "Tidy the kitchen," the system is told to gather coffee bags in one place and discard spoiled milk and orange juice cartons; it makes repeated trips around the room, tracking what has already been cleaned as its view changes with each move. In "Retrieve the medicine," the target item starts out of view inside a drawer; the system recalls the drawer's location from its stored keyframes, uses its right hand to hook and pull the drawer open, and its left hand to discard a carton at the same time — illustrating that it can select between the two arms depending on which side of the body a target sits on.

Usability Analysis

The published compute footprint is modest by robotics standards: HomeBody's local skills, perception, and motion planning run on a single Razer Blade laptop with an RTX 4090 GPU, while GPT-6 Astra itself runs remotely and exchanges skill requests and results with that laptop over the network. That split — heavy reasoning in the cloud, lightweight execution locally — is a deliberate design choice the team highlights as making the setup practical to run outside a lab. Localization inside the reconstructed environment uses Super Odometry for the robot's own position and iterative closest point registration to align the SLAM map with the simulated twin, so the planner always has a shared coordinate frame linking "where the robot is" and "where things are."

Because the paper is not yet available and the code repository holds only a README, outside researchers cannot currently reproduce the results or inspect the underlying skill implementations. What is verifiable is limited to the two demonstrated tasks in one kitchen, described qualitatively on the project page — there are no published quantitative success rates, trial counts, or comparisons to a learned-VLA baseline.

Pros and Cons

Pros:

  • Removes a training step for the VLA layer entirely, letting a swappable frontier VLM drive skills directly through structured tool calls
  • Builds persistent spatial memory from the robot's own exploration data, so it can act on objects that have left its current field of view
  • Demonstrates two genuinely long-horizon tasks (multi-object cleanup, occluded-object retrieval) rather than a single scripted pick-and-place
  • Runs its local perception and control stack on a single consumer-grade laptop GPU rather than a rack-mounted server

Cons:

  • ships with no public paper and no released code on the project's own page, so the approach cannot yet be independently verified or reproduced
  • reasoning latency from the remote frontier VLM introduces pauses between skills, by the team's own account in the page's "Scope and limitations" section
  • extended operation runs into hardware endurance limits, including finger-servo overheating, that the team lists as a current constraint rather than a solved problem
  • covers only two tasks in a single kitchen, so it is not yet evidence of general-purpose performance across varied homes or object types

Outlook

HomeBody is a proof-of-concept for a specific architectural bet: that a sufficiently capable VLM can replace a learned intermediate layer if it is given persistent memory and a well-designed skill interface, rather than needing to be trained end-to-end for the task. That bet depends heavily on the underlying VLM's tool-use and spatial reasoning improving over time, since HomeBody's own scope-and-limitations section attributes its current pauses and setup cost directly to the frontier model's reasoning latency and to Real2Sim reconstruction overhead. If a paper and code release follow the "coming soon" markers, the more interesting open question is whether the five documented skills generalize to new environments and object categories without hand-authoring new ones, and whether the approach holds up outside a single research kitchen.

Conclusion

HomeBody presents a real architectural challenge to the standard VLM-VLA-controller pipeline in humanoid robotics, backed by two concrete demonstrations in an unseen kitchen. It will interest robotics researchers and VLM-tooling teams tracking how far a frontier model's own reasoning can substitute for task-specific training. It is not yet reproducible work — no paper, no code, and no quantitative results beyond two showcased tasks are public — so readers should treat the project page's claims as an early research preview rather than a benchmarked or deployable system.

Editor's Verdict

Stanford HomeBody: A Frontier VLM Skips the VLA Layer is a workable proposition that fills a clear gap, even if it doesn't fundamentally change the landscape.

The strongest case for paying attention: removes a training step for the VLA layer entirely, letting a swappable frontier VLM drive skills directly through structured tool calls. That alone raises the bar for what readers should expect in this space. Reinforcing that, builds persistent spatial memory from the robot's own exploration data, so it can act on objects that have left its current field of view — practical value rather than just headline appeal. The broader signal worth registering is straightforward: removing the learned VLA layer shifts the burden of generalization onto the frontier VLM's own tool-use and spatial reasoning, rather than onto task-specific training data. On the other side of the ledger, one constraint is real rather than a marketing footnote: ships with no public paper and no released code on the project's own page, so the approach cannot yet be independently verified or reproduced. It should factor into any serious decision. Layered on top of that, reasoning latency from the remote frontier VLM introduces pauses between skills, by the team's own account in the page's scope and limitations section — which narrows the set of teams for whom this is an obvious yes.

For ML researchers, technical leads, and readers tracking the underlying science behind new capabilities, the smart move is to track its trajectory and revisit once the rough edges are filed down. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.

Advertisement

Pros

  • Removes a training step for the VLA layer entirely, letting a swappable frontier VLM drive skills directly through structured tool calls
  • Builds persistent spatial memory from the robot's own exploration data, so it can act on objects that have left its current field of view
  • Demonstrates two genuinely long-horizon tasks (multi-object cleanup, occluded-object retrieval) rather than a single scripted pick-and-place
  • Runs its local perception and control stack on a single consumer-grade laptop GPU rather than a rack-mounted server

Cons

  • ships with no public paper and no released code on the project's own page, so the approach cannot yet be independently verified or reproduced
  • reasoning latency from the remote frontier VLM introduces pauses between skills, by the team's own account in the page's scope and limitations section
  • extended operation runs into hardware endurance limits, including finger-servo overheating, that the team lists as a current constraint rather than a solved problem
  • covers only two tasks in a single kitchen, so it is not yet evidence of general-purpose performance across varied homes or object types
Advertisement

Comments0

Key Features

1. Replaces the usual learned System 1 VLA layer with a frontier VLM (GPT-6 Astra) calling directly into a composable library of five real-robot skills: pick, place, navigate, open drawer, pick-from-drawer. 2. Builds persistent spatial memory through a two-step deploy process: Explore (iPhone video, D435i camera, LiDAR/SLAM, joint poses, Astra-chosen waypoints) then Real2Sim (Astra builds a digital twin in NVIDIA Isaac Sim). 3. Demonstrated on a Unitree G1 in a previously unseen kitchen on two tasks: multi-object cleanup ("Tidy the kitchen") and occluded-object retrieval ("Retrieve the medicine"), without environment-specific training. 4. Runs local perception, planning, and skill execution on a single Razer Blade laptop with an RTX 4090 GPU, while GPT-6 Astra reasons remotely. 5. Project page lists its own scope and limitations: Real2Sim setup time/API cost, finger-servo overheating on extended use, and VLM reasoning latency between skills.

Key Insights

  • removing the learned VLA layer shifts the burden of generalization onto the frontier VLM's own tool-use and spatial reasoning, rather than onto task-specific training data
  • persistent spatial memory built during an exploration phase is what lets the system act on objects outside its current camera view, not any single perception model
  • structured tool calls let the VLM pick a skill and target without knowing that skill's underlying implementation, which is the interface that makes skills swappable
  • compute is deliberately split between a remote frontier model and a single consumer GPU laptop for local execution, avoiding the need for on-robot high-end compute
  • two demonstrated tasks in one kitchen is a research preview, not evidence of generalization across homes, object types, or longer task chains
  • reasoning latency and finger-servo overheating are limitations the authors disclose themselves, not just criticism from outside coverage
  • code and paper are both marked as forthcoming on the project's own page, so no outside group can currently reproduce or audit the results

Was this review helpful?

Share

Twitter/X
Advertisement