IEEE Robotics and Automation Letters · 2026 · an animated walkthrough

Learning Whole-Body Humanoid Locomotion via Motion Generation and Motion Tracking

The paper in 100 seconds · narrated · sound on
Transcript

A humanoid robot climbs onto a waist-high box, and jumps back down. It vaults over hurdles, and takes the stairs. How did it learn to move like a person?

Humanoids are built for our world of stairs, boxes and low walls. People cross them with the whole body — hands and knees included. Trial and error with a score alone rarely finds such moves. And replaying a recorded human motion fails once the box is taller.

This paper splits the job, like a choreographer and a dancer. Five minutes of human motion become an hour of robot training data. The choreographer is a diffusion model: an AI that turns noise into motion. From a height scan of the ground and a goal direction, it sketches the next half second. The dancer is a tracking policy, trained in simulation, that turns each sketch into commands for 23 joints. Because sketches can glitch, the dancer then rehearses with the choreographer — learning to keep the gist and drop unsafe moves.

In simulation, replaying one recording succeeded 23% of the time on an 80 cm box. Sketching on the fly: 96%. On the real robot everything runs on board — and it even walks around a box when it isn't set up to climb.

New skills come from more motion examples, not a new training pipeline. The catch: it needs a good LiDAR map. Learning Whole-Body Humanoid Locomotion via Motion Generation and Motion Tracking.

Robot footage: Zhang et al., Robotic Systems Lab, ETH Zurich, from the project page (CC BY-SA 4.0). Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

A humanoid robot learned to climb boxes, vault hurdles and take stairs the way people do — hands and knees included — by pairing an AI that sketches the next half second of human-like movement from the ground it sees with a controller that carries that sketch out safely.

  1. Why this matters

    Humanoids are meant to work in places built for people: stairs, crates, low walls. A person crossing a waist-high box doesn't just step — they put their hands down, bring a knee up, and hop off with their hands softening the landing. To go where people go, a robot needs the same whole-body moves, chosen from what it sees.

  2. What makes it hard

    A humanoid has many joints and balances on two small feet. Teaching it purely by trial and error with a score is like teaching someone to climb by only saying "warmer, colder": the search is enormous, and robots tend to end up with leg-only strategies while the arms just hang along.

  3. What people did before

    Copying recorded human motion gives graceful whole-body moves, but only replays that one choreography: make the box taller and it fails. Training one expert per obstacle and merging them works, but needs careful hand-design and has not scaled to many skills. AI models that invent motion were shown mostly on animated characters, and can be slow and glitchy.

  4. What this paper does

    It splits the job in two, like a choreographer and a dancer. The choreographer — an AI trained on about an hour of human-derived motion — looks at the terrain and the goal direction and sketches the next half second of movement. The dancer — a controller trained in simulation — follows the sketch with the real body. Then the dancer rehearses with the choreographer until it can follow imperfect sketches safely.

  5. What they showed

    A Unitree G1 humanoid, with all computing on board, climbed onto a 75 cm box and jumped down three different ways, vaulted a row of hurdles, took stairs, and crossed a course mixing all three. In simulation, replaying a fixed recording succeeded only 23% of the time on an 80 cm box; sketching on the fly succeeded 96%.

  6. Why it's a step forward

    New obstacle skills can come from adding motion examples instead of engineering a new training pipeline, and the choreographer is fast enough to run on the robot. One honest limit: the robot relies on a LiDAR height map, and when that map is poor, performance drops sharply.

Words used below
Reinforcement learning (RL)
learning by trial and error in simulation, guided by a score (the reward).
Motion tracking
a controller trained to copy a given reference motion with the robot's body.
Diffusion model
an AI that starts from random noise and cleans it, step by step, into a motion.
Retargeting
converting a person's movement to the robot's body proportions and joints.
Height scan
a grid of ground heights around the robot, read from its LiDAR map.
Fine-tuning
extra training of an already trained network under new, harder conditions.
1 / 8
human motion data motion generator motion tracker
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper (CC BY-SA 4.0); the text is a plain-language walkthrough.

Scene 1

Humanoids should go where people go

In short: to cross a box, a hurdle or a staircase the way people do, a humanoid has to use its arms and knees too — and pick the right move from what it sees.

Legged robots trained in simulation with reinforcement learning (trial and error guided by a score) have become remarkably agile, especially four-legged ones. Humanoids are harder: they have more joints, balance on two feet, and their best moves over big obstacles involve the whole body. Think of how you climb onto a high box: hands on the top, a knee up, then stand. Jumping down, you let your hands take part of the landing.

The paper's goal is a humanoid that does exactly this, on a real robot: walk toward a goal, see the obstacles in its path with its onboard LiDAR, and switch between whole-body skills — climbing, jumping down, vaulting, stairs — as the terrain demands.

Real run. The Unitree G1 climbs onto a box made of two stacked pallets with hands and a knee, then jumps back down. Footage: Zhang et al., Robotic Systems Lab, ETH Zurich (project page, CC BY-SA 4.0).
Paper Fig. 1. The behaviours shown on hardware: box climbing, box descent, vaulting, and a mixed sequence of obstacles.
Scene 2

Three routes, and where each one stops

In short: earlier methods either never discover whole-body moves, can only replay one fixed recording, or need a hand-built pipeline for every new skill.

1 · Learning from a score only

Training a humanoid controller purely by reward shaping — writing a score for "go forward, don't fall, move smoothly" — means searching a huge space of possible motions. Without guidance, the paper notes, controllers often converge to strategies dominated by the lower body, with little coordination between arms and legs. That is fine for walking, not for climbing a 75 cm box.

2 · Copying a recording

Motion tracking (the approach behind DeepMimic and many recent humanoid systems) trains a controller to copy reference motions from people. It transfers beautifully coordinated skills — but a tracker only replays choreography. It has no way to change a climb recorded for a 50 cm box into one for an 80 cm box.

Animated. The box grows while the recording (dashed) stays made for 50 cm. The success rates are the paper's (Table II, head-on box): replaying the fixed recording falls to 0.230 at 80 cm; the full system keeps 0.962.

3 · Many experts, merged

Another route trains one expert policy per terrain and then distills them into a single student that picks the right skill from what it sees (as in the concurrent works cited as APEX and PHP). These show promising whole-body traversal, but rely on carefully designed distillation: which expert teaches where, and on what mix of terrain. According to the authors, such methods have not yet shown strong scaling to many skills or to new terrain.

And AI that invents motion?

Generative models trained on large motion collections can compose skills naturally. But their raw output can contain artifacts — feet sliding, sudden jumps between frames — that make a real robot fall. Earlier work paired a motion generator with a tracker (CLoSD, RobotMDM, and PARC, whose design this paper follows), mainly for simulated characters, where slow inference and real-world disturbances are not an issue. This paper brings that idea to a real humanoid.

Scene 3

Five minutes of people, an hour of robot motion

In short: the team recorded one example of each skill, turned it into physically possible robot motion, then multiplied it by resizing the obstacles.

The starting point is small: about 5 minutes of human motion, from videos the team filmed themselves and from public motion-capture collections (AMASS and data from OmniRetarget). There is one representative motion per skill: climbing onto a 50 cm box, vaulting a 35 cm hurdle, jumping down from a 50 cm box, stairs with steps of about 20 cm, plus walking and turning in all directions.

  1. Video → 3-D human pose. For the filmed clips, a tool called GVHMR reconstructs the person's 3-D motion.
  2. Human → robot (retargeting). An inverse-kinematics solver that respects contacts converts the motion to the G1's body. The obstacles are then placed by hand to match where the person touched them.
  3. Physics check. Rather than trust these trajectories directly, a DeepMimic-style tracking policy learns to perform them in simulation, and the motion it actually produces becomes the training data — so everything in the dataset is physically possible.
  4. Augmentation. The obstacles are scaled and small random boxes are added along the path; the motion is re-optimized to fit (penalizing going through the terrain and jerky movement), then passed through the physics check again.

The result is about one hour of motion: box climbs and jump-downs on 35–75 cm boxes, vaults over 25–45 cm hurdles, stairs with 15–20 cm steps, and walking among small boxes.

Animated. One skill, re-fitted to five box heights. Orange rings mark the hands touching the box. Motion is illustrative; the 35–75 cm range is the paper's.
Paper Fig. 2. The whole training pipeline. Left: from videos and motion capture to a curated robot motion dataset. Middle: the generator and tracker are first trained separately on that dataset. Right: the tracker is fine-tuned while the frozen generator feeds it references on stairs, hurdles, boxes and pyramid steps.
Scene 4

The choreographer: a motion generator

In short: an AI looks at the ground around the robot and the direction it should go, and sketches the next half second of whole-body motion.

The motion generator is a diffusion model, the same family of AI that generates images from noise, here following the MDM architecture as adapted by PARC. It starts from random numbers and removes noise step by step until a plausible motion appears. What it produces is guided by three inputs:

  • the target heading — which way the robot should go (computed from a goal position);
  • a height scan — the terrain height around the robot, sampled on a 17 × 11 grid;
  • the last two frames of the robot's own motion.

Its output covers the next 0.5 s in 25 frames: where the body's root is and how it is oriented, all 23 joint angles, and the positions of the body's links. Besides matching the training motions, it is trained with extra geometric penalties (on velocities, consistency between joints and body positions, and going through the terrain). The height scans and past frames are perturbed during training so the generator copes with noisy robot sensing.

Animated. Two cleaning steps turn noise into a sketch: here, the path of one hand over the next half second as the robot reaches for the box. On the robot the generator uses only 2 denoising steps. Noise levels are illustrative.

One important detail: when the generator runs alongside the robot, it is conditioned on the robot's actual last two frames, not on its own previous sketch. That closes the loop — if the robot drifts from the plan, the next sketch starts from where the robot really is.

Scene 5

The performer: a motion tracker

In short: a second network, trained by trial and error in simulation, turns the sketch into motor commands while keeping the robot balanced.

The motion tracker is a single policy trained with RL (the PPO algorithm, in the IsaacLab simulator, on one RTX 4090 GPU), first on the curated dataset alone. Each moment it sees:

  • the reference: target body velocity, 23 joint angles and speeds, and 9 key body positions;
  • its own body sense over five recent frames: rotation rate, direction of gravity, joint angles and speeds, and its previous action;
  • the 17 × 11 terrain height scan.

It outputs 23 target joint angles. The authors note a prior study found terrain input not strictly needed for accurate demonstrations, but giving the tracker the height scan from the start helps the later fine-tuning stage.

What it is rewarded for (paper Table I)

GroupTermsIdea
Imitationbase pose, base velocity, joint angles, joint speeds, 9 key body positions (relative to the body and in the world)look like the reference
Regularizationaction changes (1st and 2nd order), joint angle / speed / torque limits, torques, body accelerationsmove smoothly, respect the hardware
Task (fine-tuning only)heading tracking: 1.5 × alignment of body velocity with the target directiongo where you're told

What this means: the tracker is paid mostly for resemblance to the sketch; only in the last stage does it also get paid for heading the right way.

Animated. Imitation rewards shrink quickly as the robot drifts from the reference. Shown: the base-position term exp(−4·gap²) from Table I, with the orientation part left out.
Scene 6

Rehearsal with the choreographer

In short: the tracker is trained further while the frozen generator feeds it live sketches, so it learns to follow imperfect plans and to ignore unsafe ones.

Simply plugging the two trained networks together is not enough. The tracker learned from smooth, physically filtered recordings; the generator's live sketches, built from noisy sensing, can contain jumps between frames and other artifacts. That mismatch makes a naive combination fragile.

So there is a third stage, fine-tuning. The generator is frozen and runs in the loop in simulation, conditioned on the robot's real last two frames, with only 2 denoising steps and extra noise injected into both the generator and the tracker's inputs. Goal directions are randomized, a heading reward is added, and the terrain gets harder and more varied: stairs with 15–25 cm steps, rows of 25–55 cm hurdles, and several boxes and pyramid steps of 30–85 cm with varying widths, yaw and pitch.

Animated. The fine-tuned tracker acts as a "motion filter": it keeps the sketch's pattern but not its jumps. Curves are illustrative.

After fine-tuning, the authors describe the tracker as a motion filter: it follows the generator's patterns and uses the height scan to suppress unsafe parts. Behaviours appear that were not in the recorded data, such as climbing a box from its corner, turning around on top of a box before jumping down, and vaulting several hurdles in a row.

On the real robot they also saw a kind of local navigation: when the robot was not in a good position to climb, the tracker partly overrode the sketch and walked around the box instead, still reaching the goal.

Real run. Instead of climbing, the robot walks around the box to reach the goal behind it (paper Fig. 3C). Footage: Zhang et al., Robotic Systems Lab, ETH Zurich (project page, CC BY-SA 4.0).
Paper Fig. 4. Same generator, tracker with (blue) and without (orange) fine-tuning, 500 simulated robots per task. Fine-tuning helps on every task, and the gap widens as boxes and steps get taller.

What this means: the rehearsal stage is not a small polish. Without it the combined system fails more and more often as obstacles grow; with it, it stays reliable across most of the tested heights. (The paper reports this as a chart only, so no exact values are quoted here.)

Scene 7

Sketching on the fly beats replaying

In short: when obstacles differ from the recording, replaying it breaks down, while the generator-plus-tracker system keeps succeeding.

To test whether live motion generation matters, the authors compare two versions in simulation, using the same fine-tuned tracker:

  • Tracker only replays one fixed reference: a climb recorded on a 50 cm box, a vault over a 35 cm high, 20 cm wide hurdle, or stairs with 20 cm high, 30 cm deep steps. Success = reaching the reference's final position.
  • Tracker + generator is the full system: the generator sketches references online from the goal direction and the terrain. Success = reaching a goal placed at that same position.

The terrain is then changed: taller or shorter obstacles, or rotated (yaw). Each setting uses 500 robots with random starting poses.

SettingHeightYawTracker onlyTracker + gen.
Box climbing (reference: 50 cm)
40 cm0°0.9821.000
50 cm−30° / 30°0.988 / 1.0000.980 / 0.982
50 cm−15° / 15°0.974 / 0.9701.000 / 0.998
50 cm0°0.9981.000
60 cm0°0.9840.996
70 cm0°0.6060.966
80 cm0°0.2300.962
Mean ± SD0.859 ± 0.2520.987 ± 0.014
Vaulting (reference: 35 cm)
25 cm0°0.9781.000
35 cm−40° / 40°0.808 / 0.4921.000 / 1.000
35 cm−20° / 20°0.922 / 0.9541.000 / 1.000
35 cm0°0.9881.000
45 cm0°0.9501.000
55 cm0°0.3480.920
Mean ± SD0.805 ± 0.2310.990 ± 0.026
Ascending stairs (reference: 20 cm steps)
15 cm0°1.0001.000
20 cm−30° / 30°0.948 / 0.9161.000 / 0.994
20 cm−15° / 15°0.970 / 0.9901.000 / 0.996
20 cm0°0.9801.000
25 cm0°0.1140.986
Mean ± SD0.845 ± 0.3000.997 ± 0.005

What this means: the fine-tuned tracker alone tolerates small changes (a slightly different height, a rotated obstacle), but once an obstacle needs a different timing or style of movement — an 80 cm box, 25 cm steps, a 55 cm hurdle — only live sketching keeps it working. At 50 cm with ±30° yaw, replay is marginally better (0.988 / 1.000 vs 0.980 / 0.982).

Real run. Vaulting a sequence of hurdles, typically crossing each directly rather than stepping onto it (paper Fig. 3B). Footage: Zhang et al., Robotic Systems Lab, ETH Zurich (project page, CC BY-SA 4.0).
Scene 8

On the real robot

In short: everything — sensing, mapping, sketching and control — runs on the robot, with a fresh sketch four times a second.

  • Where am I? DLIO, a LiDAR-inertial odometry method, uses the G1's Livox MID-360 LiDAR and its IMU to estimate the robot's pose.
  • What's the ground like? Elevation Mapping CuPy builds a terrain height map from the LiDAR and that pose; the height scans are read from it.
  • A wobbly head. The G1's neck joint is passive, so aggressive moves swing the head, which carries the LiDAR, and spoil the pose estimate. The team combines the torso IMU with the LiDAR's IMU to estimate the head's pitch and correct for it.
  • Speed. The generator runs with TensorRT on a Jetson Thor mounted on the robot's back, taking about 0.02 s per sketch. The tracker and everything else run on the G1's built-in Jetson Orin.
  • Re-planning. A sketch covers 0.5 s. In fine-tuning a new one was made every 0.5 s; on the robot a new one is made every 0.25 s (a "receding horizon"), so the robot reacts faster to what it sees.
Animated. Each teal bar is one sketch; the dark part is what the robot actually follows before the next sketch replaces it. Timings are the paper's; the animation runs at about half speed.

What the hardware tests showed

The robot climbed a 75 cm box and jumped down in three ways (straight, after a 90° turn on top, and from the corner), using knees and hands for support going up and hands for cushioning coming down — the style of the human data. It vaulted hurdles of different heights, climbed and descended stairs, and crossed a mixed course of a vaulting obstacle, stairs and a box, switching styles on its own — including combinations such as stairs right after a jump-down that it never saw together during fine-tuning.

Real run. The mixed course, filmed wide: the robot changes movement style as the terrain changes. Footage: Zhang et al., Robotic Systems Lab, ETH Zurich (project page, CC BY-SA 4.0).
Paper Fig. 3. Hardware results: (A) three ways onto and off a 75 cm box; (B) stairs and a row of hurdles; (C) walking around a box to reach a goal; (D) the mixed course.
Wrap-up

Limits and what's next

In short: pairing a motion generator with a fine-tuned tracker gives a real humanoid human-style, terrain-aware skills without a hand-built pipeline per skill — but it is only as good as its terrain map.

  • Map dependence. The system relies on a LiDAR elevation map. When mapping degrades because of sensor noise, the authors report that locomotion can deteriorate substantially. They point to learned ("neural") mapping or a belief encoder that falls back on the body's own sensing as remedies.
  • Beyond walking. Extending the approach to loco-manipulation (moving while handling objects) and to harder outdoor environments is left for future work.
  • Scope of the numbers. The quantitative comparisons are in simulation (500 robots per setting); the hardware experiments are demonstrations, shown in Fig. 3 and the video, without success rates.