RT-2 · OpenVLA · π₀ · π₀.₅
From discrete action tokens to flow matching, then to long-horizon cleanup in unseen homes. Most later humanoid VLAs branch from this line.
Notes on other people’s work, mainly humanoid VLA and world models, 2024–2026. These are not my papers.
From discrete action tokens to flow matching, then to long-horizon cleanup in unseen homes. Most later humanoid VLAs branch from this line.
Slow VLM plus a fast action expert. Helix is a company blog, not a paper. GR00T is the open, humanoid-targeted stack.
The interface moves from hand-written end-effector commands to learned latents. Large-space loco-manipulation only became its own paper in late 2025.
World models first became a data factory, then in 2026 some work treats the world model as the policy. MotionWAM is one of the few WAMs aimed at real-time humanoid loco-manipulation.
Fine-tune a video world model into a policy without changing the architecture. HumanoidArena scores the VLA–tracker interface, not just tabletop success.
A public humanoid-fleet dataset. ViLLA learns a latent planner from action-free video; WholeBodyVLA still uses that idea.