A humanoid robot learned to climb boxes, vault hurdles and take stairs the way people do — hands and knees included — by pairing an AI that sketches the next half second of human-like movement from the ground it sees with a controller that carries that sketch out safely.
Why this matters
Humanoids are meant to work in places built for people: stairs, crates, low walls. A person crossing a waist-high box doesn't just step — they put their hands down, bring a knee up, and hop off with their hands softening the landing. To go where people go, a robot needs the same whole-body moves, chosen from what it sees.
What makes it hard
A humanoid has many joints and balances on two small feet. Teaching it purely by trial and error with a score is like teaching someone to climb by only saying "warmer, colder": the search is enormous, and robots tend to end up with leg-only strategies while the arms just hang along.
What people did before
Copying recorded human motion gives graceful whole-body moves, but only replays that one choreography: make the box taller and it fails. Training one expert per obstacle and merging them works, but needs careful hand-design and has not scaled to many skills. AI models that invent motion were shown mostly on animated characters, and can be slow and glitchy.
What this paper does
It splits the job in two, like a choreographer and a dancer. The choreographer — an AI trained on about an hour of human-derived motion — looks at the terrain and the goal direction and sketches the next half second of movement. The dancer — a controller trained in simulation — follows the sketch with the real body. Then the dancer rehearses with the choreographer until it can follow imperfect sketches safely.
What they showed
A Unitree G1 humanoid, with all computing on board, climbed onto a 75 cm box and jumped down three different ways, vaulted a row of hurdles, took stairs, and crossed a course mixing all three. In simulation, replaying a fixed recording succeeded only 23% of the time on an 80 cm box; sketching on the fly succeeded 96%.
Why it's a step forward
New obstacle skills can come from adding motion examples instead of engineering a new training pipeline, and the choreographer is fast enough to run on the robot. One honest limit: the robot relies on a LiDAR height map, and when that map is poor, performance drops sharply.
- Reinforcement learning (RL)
- learning by trial and error in simulation, guided by a score (the reward).
- Motion tracking
- a controller trained to copy a given reference motion with the robot's body.
- Diffusion model
- an AI that starts from random noise and cleans it, step by step, into a motion.
- Retargeting
- converting a person's movement to the robot's body proportions and joints.
- Height scan
- a grid of ground heights around the robot, read from its LiDAR map.
- Fine-tuning
- extra training of an already trained network under new, harder conditions.