IEEE Robotics and Automation Letters · 2022 · an animated walkthrough

Combining learning-based locomotion policy with model-based manipulation for legged mobile manipulators

The paper in 97 seconds · narrated · sound on
Transcript

A four-legged robot with an arm on its back. It trots over loose blocks, keeps its hand steady when shoved, and drags a chair. How do its legs and its arm work together?

Legs go where wheels can’t — over rubble, up steps. An arm lets the robot do the job once it’s there. But the arm is heavy. Every swing pushes and twists the body — a “wrench”. Legs that can’t see it coming react late, and the body tips.

Planning with a physics model moves arms precisely, but struggles on rough ground. Learning by trial and error walks robustly, but needs retraining for each new arm.

This paper uses both. The arm planner knows how the arm will move, so it forecasts the push up to 0.8 s ahead — and the learned legs read that forecast. In training there’s no arm at all: the legs practise against invented pushes — some announced, some hidden.

Warned of a pull, the robot leans against it before it lands — and holds 150 N without stepping, about twice as much as with no push information. With a swinging 3 kg tool, the body tilts about 3° instead of 10° — and it stays level with arms it never trained with.

On the real robot it walks rough ground while the arm works — a step toward legged robots that work with their hands where wheels can’t go. Learned legs, planned arm, joined by a forecast.

Footage: Robotic Systems Lab, ETH Zurich (the paper's video). Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

A four-legged robot with an arm on its back can go where wheels can’t and do useful work there — but every swing of the arm shoves its body. This paper lets the arm’s planner warn the legs, up to 0.8 seconds ahead, how hard the arm is about to push, so the legs can brace in time.

  1. Why this matters

    Legged robots can cross rubble, climb stairs and walk where wheels and tracks can’t. Put an arm on one and it can also do things there — open doors, carry and pull loads — in places too harsh or risky for people.

  2. What makes it hard

    The arm is heavy. Every move pushes and twists the body the legs are balancing on uneven ground, while the hand still has to reach precise points. It is like carrying a tray of full glasses along a rocky trail: your legs need to know what your arms are about to do.

  3. What people did before

    Planning with a physics model moves an arm precisely and worked for legged robots with arms on flat ground, but rough ground — slips, surprise contacts — makes planning too slow. Learning by trial and error in simulation gives very robust walking, but a policy is tied to the robot it trained on: a new arm means training again.

  4. What this paper does

    It splits the job: learned legs, planned arm. The glue is a forecast. The arm planner already knows how the arm will move, so it can tell the legs what push is coming. The legs learn to use such forecasts in simulation against invented pushes — no arm needed during training.

  5. What they showed

    With a swinging 3 kg tool, the body tilted 2.9° on average instead of about 10° without push information. It worked with arms it never trained with, and held a 150 N sideways pull (the weight of about 15 kg) without stepping — twice what a policy with no push information held. On the real robot, it walked over wooden blocks and up a step while carrying and using its arm.

  6. Why it’s a step forward

    Precise arm control and robust learned walking, joined by one simple signal — and arms can be swapped without retraining. The authors describe it as the first published legged mobile manipulation on rough terrain. The legs still only react to the arm; letting them ask the arm for help is left for future work.

Words used below
Wrench
a push plus a twist on the body: three forces and three torques.
MPC
model predictive control: plan the next seconds with a physics model, re-plan constantly.
RL policy
a neural-network controller learned by trial and error in simulation.
Inverse dynamics
working out the forces a planned motion needs — and puts on the body.
Teacher / student
a policy that sees simulator secrets, then one that copies it using real sensors.
1 / 8
Arm planner (MPC) Push forecast (wrench) Learned leg policy (ours)
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough.

Scene 1

Why a legged robot with an arm?

In short: legs get a robot to hard-to-reach places; an arm lets it do real work once it is there.

Legged robots can cross rough, unstructured ground that wheeled or tracked robots can’t reach, and use buildings made for people — stairs, doors. They can also carry task-specific equipment into places too harsh for people. Mounting an arm extends what they can do there much further.

The robot in this paper is ALMA: an ANYmal quadruped from ETH Zurich with a Dynaarm manipulator on its back. The paper tests it with three arm set-ups — the bare arm, the arm holding a brick, and a 1.8 kg wrist — and walks it over obstacles.

Paper Fig. 1. ALMA in the lab: with the bare arm, with a fixed brick, with a 1.8 kg wrist, and trotting over obstacles.
Scene 2

The arm shoves the body

In short: a heavy, moving arm keeps pushing and twisting the body, and legs that can’t see it coming only react after the fact.

Controlling a legged robot with an arm is, in the paper’s words, a high-dimensional, contact-rich and often non-smooth control problem — still an open problem. The arm’s weight and motion act on the body as a wrench: a push along three directions plus a twist about three axes. When the arm swings to one side, the body is pulled down on that side; when it swings back fast, the reaction pushes the other way.

Animated. A wrench has six parts: three pushes and three twists. The paper’s forecast gives all six.

A walking controller could learn to cope with pushes it can’t see — earlier work toughened policies with random shoves during training. But an arm’s push isn’t random: its planner knows it in advance. This paper asks what happens if the legs are told.

Scene 3

Two schools of control

In short: model-based planning is precise but struggles with messy ground; learning is robust but hard to transfer. The paper uses each where it is strong.

Planning with a model

Model predictive control (MPC) repeatedly solves an optimisation: given a physics model, find the motion over the next seconds that best achieves the goal while respecting limits, apply the first bit, and re-plan. It controls wheeled mobile manipulators precisely and is robust to disturbances. For legged robots, though, every choice of which foot touches the ground when multiplies the options. Simplified “template” models made online planning possible on flat ground — including earlier whole-body work on ALMA — but on rough terrain, slips and unplanned contacts make even following a hand target while walking hard.

Learning by trial and error

Reinforcement learning (RL) has produced very robust walking on rough terrain, including the lab’s own earlier work. Contacts don’t make it more expensive, and it learns to handle slips from experience. Random pushes during training make it robust to disturbances it can’t observe. The catch: a policy usually works only on the robot and task it was trained for; a different arm means retraining.

The split

Other work had already combined the two by letting a learned policy output a goal (a body velocity, a foothold) for a model-based leg controller. This paper splits the other way: learned legs and a model-based arm, trained separately, joined by the arm’s predicted wrench. Because the leg policy never sees a particular arm during training, it can work with several arms without retraining — something a single policy trained end-to-end for the whole robot could not do.

Scene 4

The forecast

In short: the arm planner turns its plan into a forecast of the push on the body, and the leg policy reads that forecast.

Paper Fig. 2. Training (green) uses a push generator; deployment (blue) swaps in the arm planner, whose plan is converted into the push sequence by inverse dynamics.

The arm planner

The MPC treats ALMA as an arm on a moving base. Its state is the base position and orientation plus the arm’s joint angles and speeds; its inputs are the base’s forward, sideways and turning speeds and the arm’s joint accelerations. The base is modelled very simply: it follows the commanded speeds, and its height, roll and pitch settle back to their targets like a spring with damping (first-order dynamics, tuned to how the trained leg policy actually behaves).

x = [base position & orientation, θ₁…θ_N, ω₁…ω_N]  // arm joint angles θ, speeds ω
u = [v_x, v_y, ω_z, α₁…α_N]  // base speeds for the legs, arm joint accelerations α

The cost rewards reaching the hand target and penalises effort, joint speeds and straying from a nominal arm pose; joint limits are hard constraints. It is solved with SLQ in the OCS2 toolbox, running at over 200 Hz so the robot never drifts far from the plan. From the planned motion, inverse dynamics (Pinocchio library) computes the arm’s joint torques and the wrench on the base along the plan.

Animated. The five forecast samples slide along the planned push. Signal shape is illustrative.

What the legs are told

The leg policy receives the wrench at 0, 0.2, 0.4, 0.6 and 0.8 s ahead (six numbers each, in the body’s frame), the base speed command, the body’s own velocity, plus its usual senses: joint states, a history of them, and a height scan of the ground. It outputs a stepping rhythm for each leg and small joint-angle corrections for the leg motors.

Which part of the forecast matters?

To check, the authors blanked parts of the forecast (replacing them with random values) and measured how much the policy’s reward dropped over an hour of simulation. Blanking the imminent samples (now and 0.2 s) hurt most. Among the later ones, the policy relied more on the 0.8 s sample than on 0.4 and 0.6 s — the authors suggest the furthest sample tells it most about the trend.

Paper Fig. 11. Reward after blanking parts of the forecast (darker = worse). Blanking everything or the first two samples hurts most.
Scene 5

Practice pushes

In short: in training there is no arm at all — the legs practise against invented but realistic pushes, some announced and some hidden.

Simulating the arm and solving its planner for thousands of training robots would be very expensive. Instead a push generator invents wrench sequences with the smoothness of real arm plans. The paper reports this drastically reduces the cost of training (done in the Raisim simulator).

The announced part: smooth random curves

At the start of each training episode, three random values are drawn for each of the six wrench parts — for now, 1 s and 2 s ahead — and a quadratic curve is fitted through them. Every 0.02 s the curve slides forward: the values at 0 and 1 s ahead come from the old curve, and a new value 2 s ahead is drawn close to the previous one and clipped to limits. A setting β, drawn per episode, controls how far that new value may move, i.e. how fast the push changes.

w₀, w₁, w₂ ~ Uniform(w_min, w_max)  // fit a quadratic through t = 0, 1, 2 s
w(t+2+dt) ~ Uniform(w(t+2) − β·w_min, w(t+2) + β·w_max), clipped  // every dt = 0.02 s
Animated. The generator in action: squares are the three anchors, dots the five samples the policy is told. Scale illustrative.

The hidden part

The planner computes the arm’s wrench as if the base stood still, but a trotting base rocks back and forth, and that adds forces proportional to the arm’s (or payload’s) inertia. To make the policy robust to this, training adds a hidden wrench: two random vectors drawn per episode, multiplied by the body’s linear and angular acceleration, plus random noise for effects such as cogging in the arm’s motors. The policy is not told this part.

Paper Fig. 3 (a, d). Left: how fast the invented force changes for a slow (β = 0.5) and fast (β = 3.0) setting. Right: one invented force sequence, with the jagged hidden part on top of the smooth one.

Teacher and student

Training follows the lab’s teacher–student recipe. A teacher is first trained by trial and error (PPO) with privileged information only a simulator has — ground friction, contact forces — plus the exact forecast. The reward follows the lab’s earlier rough-terrain walking work, with extra weight on keeping the body level and still. Then a student learns to copy the teacher’s actions using only what the real robot has, with a noisy forecast. The student reads the forecast through a small memory network (a recurrent network), and a separate memory network also learns to estimate the current external wrench from the body’s own motion.

Animated. The student never sees simulator secrets; it copies the teacher using what the robot can actually sense.
Paper Fig. 4b. The student: a wrench memory module reads the forecast with the body’s senses; a separate belief module reads the height scan and also decodes the current external wrench.

Does the forecast help learning at all? Across five training runs each, teachers given the forecast learned faster and reached a higher reward than teachers given nothing — and training against smooth announced pushes was also smoother than against intermittent shoves.

Paper Fig. 5. Teacher training reward over 5 billion simulation samples: with the forecast (blue), with observed pushes (green), with no push information (red).
Scene 6

Bracing before the pull

In short: because it knows a pull is coming, the robot leans against it in advance — and withstands about twice the pull.

In simulation, a sideways force lasting 1 s is applied to the standing robot, and the policy sees it in its forecast. Each test starts from the same pose and gait phase; the force is raised in 10 N steps until the robot has to step to keep its balance. Five trained students were tested per setting.

Paper Fig. 9. With a 100 N pull on the way, ANYmal leans away from it beforehand, stays stable while it acts, and straightens up as the forecast shows it ending.
Simulation, from the paper’s video. The sideways pull grows in steps (yellow label, 60 → 100 N in this excerpt). Left: ours, told the forecast; middle: reactive, told only the current push; right: naive, told nothing — it is the first to fall over. Footage: Robotic Systems Lab, ETH Zurich (the paper’s video).
Animated (illustrative motion). Ours starts leaning up to 0.8 s before the pull; the reactive policy only once it has started.
PolicyWhat it is toldLargest pull without stepping
Oursthe push forecast150 N
Reactiveonly the current pushours is 170 % of it
Naiveno push informationours is 208 % of it

What this means: knowing the pull in advance roughly doubles how hard the robot can be pulled before it has to take a step (150 N is about the weight of a 15 kg bag).

Why not just estimate the push from how the body moves?

The student already learns to estimate the current wrench from its own body motion — useful by itself, e.g. to estimate a carried load. So the authors tried feeding that estimate back instead of the forecast, while the arm’s shoulder swung back and forth faster and faster with a 3 kg tool.

Arm swing speed ω (rad/s)0245
Body angular-velocity error (deg/s) — lower is better
Forecast (ours)0.2340.2410.3240.323
Estimated from body motion0.2390.2440.5390.540
Average time before falling (s) — “–” means it did not fall
Forecast (ours)–––4.1
Estimated from body motion––301.9

What this means: for slow arm motion an estimate is as good as the forecast; for fast motion, a push you only notice after it happens comes too late (Table I).

Animated (illustrative delay). The same delay that is harmless at a slow swing puts the estimate out of step at a fast one.
Scene 7

A steady hand

In short: on the real robot, when the body is pushed, the arm planner moves the arm to keep the hand where it should be.

The split also helps in the other direction. With ALMA standing, a person pushed the base forward and backward while the arm planner tracked a fixed hand target. The hand deviated 57 % less than the base. The paper notes this was with the default planner settings, and that the planner’s cost and the arm configuration could be tuned to reject disturbances even better.

Paper Fig. 8. Pushing the real robot’s base while the arm planner holds the hand target.
Real run. A person shoves ALMA’s base back and forth; the arm planner moves the arm so the hand stays near its target — the experiment behind the 57 % figure. Footage: Robotic Systems Lab, ETH Zurich (the paper’s video).

In the video, ALMA also trots over loose wooden blocks and climbs a step with the arm on board, pulls a chair across the floor, and adjusts its posture promptly when the arm moves — the hardware demonstrations behind the paper’s claim of legged mobile manipulation on rough terrain.

Scene 8

Results in simulation

In short: any push information keeps the body more level; the forecast matters most when the arm moves fast; and it carries over to arms the policy never saw.

The full system runs in simulation on ALMA with a 3 kg tool, always trotting. The planner chases a hand target and sends body-speed commands to the legs. Three measures show whether the body behaves as the planner’s simple model assumes: mean body tilt, how fast the body rocks (pitch and roll speed), and how well the body follows the speed command. Each policy was trained with five random seeds.

  • Experiment 1: on rough terrain, a random hand target in a 3 × 3 m square at 0.6–1.1 m height, moved every 5 s; half an hour of simulated time. The base moves at up to 0.5 m/s.
  • Experiment 2: the target swings from side to side (y = 0.5·sin(τ/0.6) m, 0.5 m ahead, 1.1 m high) while the base holds its place — the push changes much faster.

Four students are compared: ours (forecast through a memory network), the forecast through a plain network without memory, the current push only, and no push information.

Paper Fig. 6. Left: Experiment 1 (random targets). Right: Experiment 2 (swinging target). Tilt in degrees, rocking in 0.1 rad/s, tracking error in 0.1 m/s.
Mean body tilt (°)OursForecast, no memoryCurrent push onlyNo push info
Exp. 1 · random targets2.02.12.14.6
Exp. 2 · swinging arm2.93.24.6≈10*
Exp. 2 · rocking (0.1 rad/s)3.53.25.34.7
Exp. 2 · tracking error (0.1 m/s)1.21.51.41.9

What this means: when the arm is mostly still, any push information halves the tilt; when it swings fast, knowing the push only as it happens is not enough — the forecast keeps the body level. *The label is cut off in the figure; the same setting in Fig. 7 reads 10.1°.

Why is the version without memory about as good as ours? The paper explains that no offset or scaling was added to the forecast in these tests, so a plain reading of the current forecast works just as well. The memory network’s advantage depends on how much noise is put on the forecast in training, and the authors saw it on hardware, but left that comparison out for consistency.

Arms the policy never saw

Experiment 2 was repeated with three arms: light (Dynaarm without tool), long (forearm stretched to 0.6 m and made heavier to match, no tool) and heavy (with the 3 kg tool). No retraining.

Paper Fig. 7. Tilt and tracking error for three arm set-ups, ours vs. the policy with no push information.
Simulation, from the paper’s video. The leg policy only needs the push forecast from the arm, so it runs with different arm set-ups without retraining; the red balls are the hand targets. Footage: Robotic Systems Lab, ETH Zurich (the paper’s video).
Mean body tilt (°)Light armLong armHeavy arm
Ours1.692.992.93
No push information4.339.4210.1

What this means: with bigger arms, the policy without push information tilts two to three times more; the forecast lets one policy handle all three.

A load nobody modelled

To test the hidden-push training, an extra 1 kg payload the planner doesn’t know about was put in the hand (Experiment 1 setting). A policy trained without the hidden, acceleration-dependent pushes was compared with ours.

MeasureOurs (hidden pushes in training)Without them
Mean tilt (°)3.03.1
Body rocking (0.1 rad/s)2.44.1
Speed tracking error (0.1 m/s)2.02.3

What this means: practising against pushes it isn’t told about makes the body markedly calmer when reality differs from the model (Fig. 10).

Wrap-up

Limits and what’s next

In short: the legs take the arm’s push as given; the next step is letting them negotiate it.

  • Passive legs. The leg policy only responds to the arm’s wrench; it has no say in it. The authors propose letting the base ask for a helpful wrench that the arm then tracks, which they expect to enable more demanding manoeuvres.
  • Simple body model in the planner. The planner assumes the body follows speed commands perfectly and settles like a spring. Mismatch is inevitable; running the planner above 200 Hz keeps the deviation small, and the hidden-push training makes the legs tolerant of it.
  • Estimates on the real robot. The student’s own wrench estimate is useful when things change slowly — the paper’s video shows ANYmal estimating the pull while dragging a chair — but it lags behind fast arm motions.
  • Beyond arms. The same policy could serve other jobs where a push is known in advance, such as carrying or pulling loads.