CoRL · 2025 · an animated walkthrough

Motion Priors Reimagined: Adapting Flat-Terrain Skills for Complex Quadruped Mobility

The paper in 110 seconds · narrated · sound on
Transcript

Meet ANYmal, walking like a dog, down outdoor steps, onto pallets and blocks, and around obstacles to a goal. Its gait came from dogs filmed on flat ground. So how does it manage this?

Legged robots should go where wheels can’t. Given just a goal, this one must cross rough ground and find its own way around tall obstacles. Robots learn by trial and error in simulation, guided by a score. Copying a dog’s recorded motion looks natural, but it was recorded on flat ground, so on stairs the copy trips. Learning from scratch climbs, but hops.

The fix comes in two layers. First, a motion prior, a layer of dog-like skills, learns the walk, pace and canter on flat ground, and is then frozen. On top, a second policy sees the terrain and the goal. It picks the gait and adds small residuals, corrections to each joint, like lifting a foot over a step. A penalty keeps them small. Too strong a penalty, and the robot stalls on tall steps. Just right, it climbs in the dog’s style. None, and the style drifts away. A student version then learns to do the same from the robot’s real, noisy sensors.

With noisy sensors, it reached a goal 5 m away in 75 to 95 of 100 tries: on tall stairs, steep slopes, and around obstacles. Without the prior, the same reward gives a hopping gait. With it, the robot walks like an animal, and uses less energy. On the real ANYmal robot, it climbed stairs, and crossed blocks, finding its own way to the goal.

Flat-ground animal skills, carried onto rough terrain with small learned corrections. Scroll down for the full walkthrough.

Robot footage: the authors' project page, anymalprior.github.io (Zhang, Li, Miki, Hutter · Robotic Systems Lab, ETH Zurich). Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

This paper teaches a four-legged robot to copy dog movement recorded on flat ground, then to learn small corrections, so the same natural gait carries it up stairs, over blocks and around obstacles to a goal.

  1. Why this matters

    Legged robots are built for places wheels can't go: stairs, piles of blocks, cluttered rooms. To be useful they must reach a goal on their own, and move smoothly and efficiently while they do it.

  2. What makes it hard

    Robots learn to walk by trial and error in simulation, guided by a score (the "reward"). Getting a natural gait out of that takes many hand-tuned penalty terms; without them the robot tends to learn an odd, hopping style.

  3. What people did before

    Copying animals ("motion imitation") gives natural movement with little tuning, but only on the kind of ground where the animal was filmed. Earlier fixes needed rough-terrain recordings with a carefully matched simulator, or specially optimized trotting motions.

  4. What this paper does

    Two layers. The bottom layer learns walk, pace and canter from flat-ground dog data and is then frozen. The top layer sees the terrain and the goal, picks the gait, and adds small corrections to each joint. One penalty keeps those corrections small.

  5. What they showed

    With noisy sensors in simulation, the robot reached a goal 5 m away in 75 to 95 of every 100 tries, on 25 cm stairs, 24° slopes, random boxes and around tall obstacles. Trained the same way without the dog layer, it hopped. The real ANYmal walked over stairs and blocks and around obstacles, indoors and out.

  6. Why it's a step forward

    Cheap flat-ground animal data becomes reusable for rough terrain and navigation in one controller, with far less reward tuning. One honest limit: the robot tends to settle on a single gait and bend it, rather than switch gaits.

Words used below
motion prior
a frozen skill layer that already moves like the animal
code (latent command)
16 numbers that tell the prior which gait and rhythm to run
residual
a small correction added to each joint's target
reward
the score that guides learning by trial and error
teacher / student
a policy with perfect information trains one that uses only real sensors
1 / 8
motion prior (dog skills) residual corrections full method · goal
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough.

Scene 1

Why it matters

In short: the robot gets only a goal position, and must both walk over rough ground and find its own way around things it can't climb, in a smooth, animal-like gait.

Controllers trained with reinforcement learning (trial and error in simulation, steered by a reward) already let legged robots follow a commanded speed or walk to a commanded position over difficult terrain. Position-based controllers have an extra benefit: they can dodge local obstacles by themselves, without a separate planner that hands them waypoints.

The catch, according to the authors, is that both kinds depend on meticulous reward design and still struggle to produce expressive, animal-like motion. This paper wants all of it at once: goal reaching, perceptive walking over stairs, slopes and blocks, local obstacle avoidance, and a natural gait, from one policy running on the ANYmal-D quadruped.

Paper Fig. 1. The real robot in the lab, on a step, over pallets, outdoors on gravel and in an office. The arrows show where it is heading; nobody gave it a path, only a goal.
Scene 2

What breaks

In short: copying animals gives beautiful gaits but only on the ground they were recorded on; learning from scratch works on any ground but needs lots of tuning to not look odd.

Motion imitation trains a policy to follow reference movement, from motion capture, teleoperation or trajectory optimization. It has produced agile skills like jumps and backflips that are very hard to get from a hand-written reward. Two families exist: tracking methods (in the DeepMimic line) reward the robot for matching the reference frame by frame, and adversarial motion priors (AMP) train a discriminator to judge whether movement looks like the data. The authors chose the tracking route, building on FLD (Li et al. 2024), because AMP-style training can be unstable and can collapse to a few motions when the dataset is diverse.

The problem is how narrow the copy is. A policy forced to mimic a fixed set of demonstrations overfits to how they were recorded. Put it on different ground and it meets states the data never showed (covariate shift), so it fails or loses its style. The dog data used here was recorded on flat ground only.

Where earlier fixes stop

  • Han et al. (2024) train a low-level imitation policy on animal motion and a high-level policy for tasks on rough terrain. It works well, but relies on motion data recorded on uneven terrain and on calibrating the simulator to match how that data was captured.
  • Wu et al. (2023) stay stable on complex terrain using priors from flat-ground data only, by adding an AMP-style reward. But their references come from trajectory optimization and mostly give a trotting gait.
  • Position-based locomotion (e.g. Rudin et al. 2022, Ren et al. 2025) shows that one policy can learn to navigate locally, but needs extensive reward shaping for smooth motion and was mostly shown on near-flat ground.

The gap this paper fills: take raw retargeted animal data (walk, pace, canter, not only trot), recorded on flat ground, and make it work on terrain it never saw, while also navigating.

Scene 3

The motion prior

In short: dog motion is mapped onto the robot, summarized into a short rhythmic code, and a low-level policy learns to play back any gait that code describes, on flat ground.

1 · From dog to robot

The dog motion capture (from Han et al. 2024 and the dog data of Zhang et al. 2018) includes walking, pacing and cantering at different speeds, all on flat ground. It is retargeted to ANYmal with inverse kinematics: the robot's pose is solved so that its neck, pelvis and feet positions match the dog's, following the pipeline of Peng et al. (2020).

Animated. The orange points (neck, pelvis, feet) are what the retargeting matches between dog and robot. Both figures are schematic.

2 · A code for rhythm

Next, an encoder from FLD (Li et al. 2024) is trained on the motion data. It looks at a window of 31 frames (0.62 s at 50 Hz) and describes it with a few periodic channels (four in this paper), each with a frequency, amplitude, offset and phase. It is trained so that just moving the phase forward in time, then decoding, reconstructs the frames that follow. That forces the code to capture the motion's rhythm, not the details of one clip.

Animated. The encoder sees the orange window; the dashed curve is what advancing its phase predicts next. Curves are illustrative.

3 · A policy that plays the code

The low-level policy gets the robot's body sense (base velocities, the direction of gravity in the body frame, joint positions) plus the code, and outputs 12 joint targets. It is rewarded for tracking the reference's body velocity, turning rate, joint positions and orientation, with small penalties for jerky actions, collisions, high torques and hard impacts. The authors changed two things from FLD: they skip the decoder and track the true reference motion directly, which avoids errors from imperfect reconstruction, while keeping the structured code.

reward ≈ 2·exp(−|v − vref|²) + 0.8·exp(−0.8|ω − ωref|²) + 1.4·exp(−2 Σ(q − qref)²) + 0.8·exp(−3|g − gref|²) − penalties  // v velocity, ω turning rate, q joint angles, g gravity direction (Table 1)
Paper Fig. 3. Joint angles on flat ground while the code switches canter → walk → pace → walk. Dashed is the dog reference, solid is the robot: they nearly coincide, and the switches are smooth.
Paper Fig. 10. When each foot touches the ground in the three learned gaits. Walk keeps most feet down; pace moves the two legs of one side together; canter has a three-beat rhythm.

Once trained, this policy is frozen. It is the motion prior: an interface that turns 16 numbers into animal-like movement.

Scene 4

Residuals

In short: a high-level policy, trained on rough terrain, drives the frozen prior with a code and adds small per-joint corrections where the flat-ground gait isn't enough.

Paper Fig. 2. The whole pipeline: (1) retarget dog data, (2) pre-train the encoder and low-level policy on flat ground, (3) train the high-level teacher on rough terrain with the prior frozen, (4) distill a student for the real robot.

The high-level teacher outputs two things: a 16-number code for the low-level policy, and a 12-number residual that is added to the low-level policy's joint targets. It observes the same body sense as the low level, height samples around each foot fused with a downsampled, cone-shaped LiDAR scan, "privileged" facts only a simulator knows (foot contacts, friction, external pushes), and the task: the goal position relative to the robot and a speed command.

Animated. Dashed is where the flat-ground gait would put the foot; the residual lifts it over the box edge. In the paper the residuals are added to joint targets; showing them as extra foot lift is a simplification.

What the robot is rewarded for

Two task rewards drive navigation. The reach reward pays out only near the end of each command period and only within 2 m of the goal, more the closer the robot is. Because that is sparse, a denser velocity reward pays for progress toward the goal, capped at the commanded speed, until the robot is within 15 cm. All the low-level policy's regularization penalties are kept with the same weights, and one new term penalizes the squared size of the residuals.

rreach = (1/Tr)(1 − ‖d‖/2)  if t > T − Tr and ‖d‖ < 2 m  // d: vector from robot to goal
rvel = min(vcmd, ⟨v, d⟩)  if ‖d‖ > 0.15 m  // progress toward the goal, capped
rres = wres Σi=1..12 (aresi)²  // wres = −0.1 in the main policy
Animated. The velocity reward (teal) pays while the robot makes progress; the reach reward (violet) only switches on inside 2 m and in the last stretch of the command period. Schematic timing and values.
High-level reward termWeight / form
Position tracking (reach)15 · rreach
Heading velocity5 · rvel
Joint residual−0.1 · Σ residual²
Termination (fall)−200
Collision, stand still, stand pose, torques, action rate, …kept from low level

What this means: the task itself is described by just two terms; the only new knob for motion quality is the residual penalty. The rest are carried over unchanged from the low level (paper Table 2).

Training runs 4096 simulated robots in parallel in Isaac Gym with PPO on a single RTX 4090, on a mix of boxes, stairs, rugged slopes and high obstacles (paper Fig. 9).

Paper Fig. 9. The simulated training world for the high-level policy: boxes, stairs, rugged slopes and tall obstacles, with blue goal markers.
Scene 5

The penalty dial

In short: how strongly the corrections are penalized decides the balance between keeping the dog's style and adapting to hard ground, and the corrections grow exactly where the ground gets hard.

Corrections grow with difficulty

The authors walked the policy across pyramid stairs and rugged pyramid slopes with five difficulty levels each (steps from 5 to 25 cm, slopes from 0 to 24°). The spread (variance) of the residuals rises with difficulty in all joints. The hip joints of the front legs show peaks that grow step after step, which the authors link to changes in body orientation and step height that the flat-ground gait alone cannot handle.

Animated. As the levels get taller, the corrections (teal trace) get bigger. The trend matches the paper; the values are illustrative.
Paper Fig. 5. Residual variance per joint (columns) at each difficulty level (rows). Darker means the corrections are doing more; it darkens toward the tall steps and steep slopes, most in the hip and knee joints.
Paper Fig. 11. The test courses: (a) pyramid stairs and (b) rugged pyramid slopes. The goal moves from easy to hard levels; the red line is one of the robot's paths.
Paper Fig. 4. The front-left and front-right hip residuals over time while the goal moves from easy to hard levels (terrain in paper Fig. 11). The peaks rise level by level.

Turning the dial

Six high-level teachers were trained with residual penalty weights of 0, −0.001, −0.1, −1, −10 and −100. The paper's findings:

  • All six reach the goal reliably on near-flat ground.
  • A penalty stronger than −10 makes the robot much less able to handle higher steps and steeper slopes. At −100 it fails on most tall stairs.
  • A penalty weaker than −0.001 lets the residuals grow without bound. Success stays high, but the gait no longer follows the prior's style, even on flat ground.
Paper Fig. 6a. Goal-reaching success on stairs, going up (left) and down (right), for the six penalty weights (legend in 6b). The strongest penalties fall off as the steps get taller.
Paper Fig. 6b. The same on rugged slopes. 100 trials per terrain, random start poses. Weights of −0.1 to 0 stay high throughout.
Paper Fig. 7. How far the joints stray from the dog reference during five seconds of walking on flat ground. The weaker the penalty, the further the gait drifts from the animal's.

What this means: instead of tuning many reward weights for good-looking motion, there is essentially one knob. Too tight and the robot can't adapt; too loose and it stops moving like the animal. The main policy uses −0.1.

Scene 6

Teacher → student

In short: the high-level policy learns with perfect information in simulation, then a student with only the robot's real, noisy senses learns to copy it.

Perfect terrain and hidden physical facts make the teacher learn fast, but a real robot has neither. So a student policy is trained by supervised learning to reproduce the teacher's outputs (the code and the residuals) from noisy body sense and noisy terrain scans only. Its training signal is simply the difference between its actions and the teacher's.

The student contains a recurrent belief encoder: a small memory network (a gated recurrent unit) that, over time, estimates the hidden facts and a cleaner version of the terrain from the noisy inputs. Its design follows Miki et al. (2022), earlier work from the same lab. The student is trained with truncated backpropagation through time over 15 steps.

What this means: the motion prior and the residual idea don't need to change for the real robot; the teacher–student step is a standard recipe for crossing the gap between simulation and reality.

Scene 7

Results

In short: with noisy sensors, the student reaches a goal 5 m away in 75 to 95 of 100 tries across seven terrains, and the real robot does the same kind of thing in the lab and outdoors.

Terrain (simulation, noisy sensors)Goal reached
0.25 m stairs, up84 / 100
0.25 m stairs, down85 / 100
24° slope, up90 / 100
24° slope, down95 / 100
Random boxes81 / 100
Random boxes with high obstacles75 / 100
Flat ground with high obstacles90 / 100

What this means: the same flat-ground dog gait, with learned corrections, gets up 25 cm stairs (taller than a typical 17–18 cm house stair) more than four times in five, and the hardest case is the one that combines rough footing with route-finding.

Each terrain was tested 100 times with the goal 5 m from the start in a random direction (paper Table 8).

On the real robot

The student runs on ANYmal-D. Given a fixed goal, it crossed random steps and stairs and walked around obstacles in the lab, and navigated outdoors, in a smooth animal-style gait and without any waypoints for going around obstacles. The paper reports these hardware runs qualitatively, with photos and videos rather than success counts.

Paper Fig. 8. Hardware sequences toward a goal (blue pin): (a) over blocks and pallets in the lab, (b) around an obstacle, (c) outdoors around a raised deck.
Real run. ANYmal-D climbing an indoor staircase with the learned policy. 7 s loop, muted. Footage: Zhang, Li, Miki, Hutter · Robotic Systems Lab, ETH Zurich (project page).
Real run. Local navigation in the lab: the red pin marks the goal, and the robot steps over blocks and pallets to get there while a person rearranges obstacles. 8 s loop, slightly zoomed, muted. Footage: Zhang, Li, Miki, Hutter · Robotic Systems Lab, ETH Zurich (project page).
Scene 8

Vs. from scratch

In short: under the same reward, a policy without the prior hops; the one with the prior walks like an animal and spends less energy, without ever being told to save energy.

The baseline has the same observations (minus the code) and outputs 12 joint targets directly, trained under the same rewards with no extra tuning or exploration tricks. It can cross challenging terrain, but it adopts a jumping gait. The authors note that extra penalties, for example on body acceleration or vertical speed, could make it more natural, but that is exactly the tuning the prior is meant to save.

Simulation. Left, the baseline trained from scratch with the same rewards; right, the policy with the motion prior, on the same block terrain. 8 s loop, muted. Footage: Zhang, Li, Miki, Hutter · Robotic Systems Lab, ETH Zurich (project page).

Energy

As a further check, the robots walked forward on flat ground following a moving goal at a range of speeds, and the cost of transport was measured (energy spent per unit weight per unit distance, lower is better). Both policies with the motion prior used less energy than the baseline at every speed tested, and removing the residual penalty raised the cost slightly. The authors' reading: the prior naturally avoids jumping gaits, but further research is needed to compare this with adding energy terms to the reward directly.

Paper Fig. 12. Cost of transport against forward speed, averaged over 20 s of walking. The from-scratch baseline (orange) is highest at every speed; the prior with the −0.1 penalty (blue) is lowest.

Five reward terms are enough

In an extra experiment, a high-level policy was trained on the same prior with only five reward terms: reach, velocity, residual penalty (−0.5), collision and termination. It still walked over varied terrain and navigated around obstacles in an animal-like gait. Interestingly, it chose to canter rather than walk, and kept the canter on hard terrain with the help of the residuals.

What this means: the natural style comes mostly from the prior, not from the reward, so the reward can stay short.

Limits

Limits and what's next

In short: the robot tends to stick with one gait, and the terrains are still a limited set.

  • One gait, bent to fit. The high-level policy can suffer from mode collapse: it habitually uses a single low-level gait and relies on residuals to adapt, instead of switching between walk, pace and canter. Exploration strategies or rewards for diversity might help.
  • Limited terrain and skills. No gaps, stepping stones or overhanging obstacles yet, and the prior contains only walk, pace and canter. Adding skills like jumping or crawling is future work.
  • A scalable direction. The authors suggest training with randomly sampled codes for different motions while commanding a fixed motion at deployment, as a way to adapt many skills at once.