IEEE ICRA · 2024 · an animated walkthrough

Learning Risk-Aware Quadrupedal Locomotion using Distributional Reinforcement Learning

The paper in 119 seconds · narrated · sound on
Transcript

One walking policy, with a dial for risk. On the real robot, ANYmal, set to bold, it goes for the high step; set to careful, it holds back. So how do you teach a robot to weigh risk?

Legged robots can go where wheels can’t: caves, forests, maybe other planets. Out there, a single fall can break the robot, or end the mission.

But learned controllers are trained on averages. A climb that usually works, but sometimes ends in a fall, can look as good as a safe detour. The usual fix, hand-tuning the reward, the score it learns from, is slow and fragile.

This paper changes how the robot learns, not what it’s rewarded for: its critic, the part that predicts how things will go, learns the whole spread of outcomes. Then one number, the risk dial, decides which outcomes count most: the worst ones for a careful robot, the best ones for a bold one. The dial is also an input to the walking policy, so one trained policy covers every setting, and it can be turned while the robot walks.

In simulation, the bold setting climbs a 55 cm barrier. The careful one refuses. On steep stairs, the bold one usually falls, while the careful one walks down safely. Set to neutral, it walks as well as the standard method: an average score of 27.6, against 27.9.

On the real ANYmal, at a 27 cm step, bold mode got up in 62.5 % of tries. At 41 cm, careful mode refused 85.7 % of the time. In the lab, the bold setting tries the step and slips back; the careful one backs away.

Risk becomes a knob: careful for a newcomer, full skill when the mission needs it, and no reward retuning.

Footage: Robotic Systems Lab, ETH Zurich (the paper’s video). Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

The authors trained a four-legged robot to foresee the whole range of good and bad outcomes of its moves, not just the average, so that one dial can tell it, even while it walks, how much risk to take.

  1. Why this matters

    Legged robots can go where wheels can’t: cave systems, forests, maybe other planets. There, one fall can break the robot or end the mission. The robot has to weigh risk, not only reach its goal.

  2. What makes it hard

    Learned walking controllers are trained to maximise their average score. A move that usually works but sometimes ends in a crash can look just as good as a safe one. It’s like choosing a commute by its average time and never noticing the route that sometimes takes two hours.

  3. What people did before

    Controllers trained by reinforcement learning already walk robustly over rough terrain, but they don’t model risk. To make them careful, engineers edit the reward: slow down here, avoid that. Reward tuning is already slow, and its terms affect each other. An existing risk-aware learning method, DSAC, didn’t learn proper walking in the authors’ tests.

  4. What this paper does

    Change how the robot learns, not what it is rewarded for. Its critic, the part that predicts how well things will go, learns the whole spread of possible outcomes. A single number, the risk dial, decides which outcomes count most: the worst ones (careful) or the best ones (bold). The dial is also an input to the walking policy, so one policy covers every setting. They call the method DPPO.

  5. What they showed

    In simulation, bold settings walk faster and try everything, sometimes crashing; careful settings go slowly and refuse a 55 cm barrier and an 85 cm pit. Set to neutral, it learns as well as the standard method (return 27.62 vs 27.86). On the real ANYmal, careful mode refused a 41 cm step 85.7 % of the time; bold mode climbed a 27 cm step in 62.5 % of tries.

  6. Why it’s a step forward

    Risk becomes a knob for an operator or a route planner: careful for an amateur at the controls, the full range of skill when the mission needs it, and no reward retuning. One honest limit: someone still has to choose the setting, and how best to measure “risk-awareness” is an open question.

Words used below
reinforcement learning
learning by trial and error to earn a high score (reward)
return
the total score collected over one run
policy
the network that turns sensor readings into joint commands
critic
a network that predicts the return from the current situation
value distribution
the whole spread of returns the critic expects, not just their average
risk metric · β
a rule that reads the spread with a bias; β sets how careful or bold
1 / 8
risk-averse neutral · value distribution risk-seeking
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough.

Scene 1

Where a single fall can end the mission

In short: legged robots are built for dangerous places, and in dangerous places a robot that can’t weigh risk is a liability.

Walking robots can handle slopes, loose ground, stones, stairs and even gaps by choosing where to put their feet. That makes them a good fit for cave systems, forests or, one day, the surfaces of other planets. The paper opens from exactly this point: in such places, the robot’s safety comes first, because a failure can mean broken hardware or a failed mission.

Controllers learned by reinforcement learning (the robot practises in simulation and is rewarded for good walking) have made legged robots reliable on difficult terrain. The authors point out what those controllers still lack: they don’t account for risk explicitly. Nothing in them says “this step usually works, but if it fails, it fails badly”.

The cave in the animation is an illustration. The 85 cm drop matches the depth of the pit in the paper’s simulated obstacle course (Scene 6).

Scene 2

Same average, very different danger

In short: a learner that only predicts the average score can’t tell a safe move from a gamble with the same average, and the usual fix of editing the reward is slow and brittle.

Standard reinforcement learning, including PPO (Proximal Policy Optimization, the workhorse behind most learned walking controllers), trains a critic that predicts one number per situation: the expected return, meaning the average total score from here on. The policy is then pushed toward actions that beat that expectation.

An average hides the shape of the outcomes. In the animation, climbing usually earns a high score but sometimes ends in a fall; walking around always earns a middling score. Both average out the same, so an average-only learner sees no difference between them.

Animated. Outcomes pile up for the two options. The average lines land in the same place; only the full spread shows the cluster of falls. Scores are illustrative.

Why not just change the reward?

Before this work, the way to get more careful walking was to change the reward: reward slower walking on rough ground, penalise approaching obstacles. The authors argue this is undesirable. Getting a good walking controller already needs expensive, slow reward tuning, and reward terms interact. For example, large torque penalties can keep the learner from ever discovering the speed-tracking reward. Scene 7 shows what happened when they tried two reward-based versions of a “risk dial” as baselines.

Risk-aware reinforcement learning did exist. The closest work learns a risk-averse policy with an adjustable level, using a Gaussian value estimate and the CVaR metric; DSAC adds a value distribution to the Soft Actor-Critic algorithm. But in this paper’s locomotion setup, DSAC did not learn to walk properly (Scene 7).

Scene 3

A critic that predicts the whole range of outcomes

In short: instead of one average, the critic predicts a set of values that together describe how likely every outcome is, including the rare, bad ones.

This is the idea behind distributional reinforcement learning: learn the entire distribution of returns, not just its mean. It was first proposed to get more accurate value estimates, but the distribution also carries exactly what risk needs: how likely catastrophic and high-return outcomes are.

The paper’s critic uses the representation from QR-DQN. For each situation s, the network outputs N numbers θ1(s)…θN(s). Each one is a stick carrying probability 1/N, and together they form the value distribution:

Z(s) = (1/N) · Σi δ(θi(s))  // N equally likely outcomes; the network decides where they sit

How the critic is trained

A classic critic learns from targets like “reward now + discounted value later”. For a whole distribution, the paper builds multi-step targets with SR(λ), a sample-replacement scheme from earlier work (GMAC). Start with samples of the predicted distribution at the last step of a rollout. Then walk backwards in time: discount the samples, add the reward received, and swap a fraction 1−λ of them for fresh samples from the critic’s own prediction at that step. With λ = 0 this is a 1-step target; with λ = 1 it is the full N-step return. The critic is then pulled toward this target by minimising the energy distance, a way to compare two sets of samples.

Animated. SR(λ) built step by step: shift and shrink, then swap in fresh samples. λ, γ and the rewards are illustrative; the procedure is the paper’s.
Paper Fig. 1. One policy at three risk settings, told to walk up a large step. The risk-averse robot (purple) refuses, the risk-seeking one (red) climbs. Below each robot is its value distribution as distorted by the risk metric, at successive moments.
Scene 4

One number decides which outcomes count

In short: a risk metric re-weights the sticks before averaging them. Favour the worst ones and the robot becomes careful; favour the best ones and it becomes bold.

The paper follows work arguing that distortion risk metrics suit robotics. Such a metric bends the quantile levels τ (τ = 0 is the worst outcome, τ = 1 the best) with a function g, and averages the sticks with the resulting weights:

Vβ(s) = Σk ( gβ(τk) − gβ(τk−1) ) · θk(s),   τk = k/N  // g(τ) = τ gives the plain average

Wang metric

gβ(τ) = Φ( Φ−1(τ) + β )  // Φ: standard normal CDF · β = 0 neutral · β > 0 averse · β < 0 seeking

The Wang metric shifts the whole curve smoothly: with β > 0 the low (bad) outcomes get more weight, with β < 0 the high ones do. The policy was trained on β ∈ [−1.5, 1.5].

Animated. As β rises, the curve bows up and the worst tenths of outcomes (orange) take most of the weight; as β falls, the best tenths do. Exact formula; bar heights computed from it.

CVaR

gβ(τ) = min( τ/β , 1 )  // β = 1 neutral · 0 < β < 1 averse: average of the worst β share

Conditional Value at Risk (CVaR) is the expected return in the worst β share of cases. It simply ignores everything better than the cut-off. Trained on β ∈ (0, 1].

Animated. CVaR keeps only the shaded worst share and averages it. As the share shrinks, the value slides toward the falls. The distribution is illustrative.

A useful finding: behaviour generalised beyond the trained β range. Using β outside the bounds gave stable but even more risk-averse or risk-seeking behaviour.

Paper Fig. 3. The risk metric as a block: the risk sensitivity selects a distortion of the value distribution, and the mean of the distorted distribution is what the learning update uses.
Scene 5

The dial is an input, so one policy learns every setting

In short: the risk-skewed value replaces the plain value inside PPO, and because β is fed to the policy, a single trained controller can be switched between careful and bold at any time.

The policy maps the robot’s state to desired joint positions. The state holds body measurements (base velocity, gravity direction, joint positions and velocities), the previous action, samples from a local height map around the robot, the velocity command and, new here, the risk parameter β. The simulation uses the robot model from legged_gym, including learned actuator dynamics.

Learning is PPO’s usual clipped update. The only change is inside the advantage, the signal for “this action went better or worse than expected”. It is computed with truncated Generalized Advantage Estimation, using the risk-aware values:

δt = rt + γ·Vβ(st+1) − Vβ(st)  // the advantage sums these with weights (λγ)^l over the rollout

So the same reward and the same outcome can produce a positive advantage for a bold robot and a negative one for a careful robot:

Animated. Same distribution, two readings. The averse reading sits below the value before the step, so climbing gets discouraged; the seeking reading sits above it, so climbing gets reinforced. Numbers are illustrative.

Why condition on β rather than train one policy per setting? No single risk preference fits every situation, so an operator or a high-level planner should be able to change it on the spot. In training, β is drawn uniformly at random along with each command. Training ran in legged_gym on Isaac Gym, over seven terrain types with ten difficulty levels each. The rewards come from legged_gym’s locomotion task (track the command, avoid collisions and wasted energy), plus an alive reward, impact-dependent collision penalties, and joint-motion and torque penalties that increase through a curriculum.

Paper Fig. 2. The full training loop. The critic predicts a value distribution, the risk metric turns it into values for the PPO update, and the risk parameter reaches the policy as part of the operator’s command.
Scene 6

Same policy, two dial settings, very different choices

In short: driven through the same obstacle course, the bold setting attempts everything and sometimes crashes; the careful setting walks slowly, never falls on the stairs, and refuses the obstacles it can’t safely handle.

To show how an operator could use the dial, the authors drove a simulated robot through a course. The shortest route goes down steep stairs (25 cm steps), over a 55 cm barrier, up a ramp and down into an 85 cm deep pit. Their observations:

  • Risk-seeking generally walked faster and tried to follow every command. It usually got over the barrier (c), occasionally getting stuck on it or falling off the far side. On the stairs (b) it sometimes made it down but fell in most cases. It stepped into the pit (e) every time and often crashed.
  • Risk-averse walked slowly but safely. It came down the stairs on every attempt, refused to climb the barrier (so an operator would route it around, path d), and refused to step into the pit (f).
  • On the incline (a), risk-seeking can be chosen simply to walk faster.
Paper Fig. 4. The obstacle course. Colored paths show which risk setting suits each part: seeking for the incline (a), barrier (c) and pit (e); averse for the stairs (b), with detour (d) around the barrier and a refusal at the pit (f).
Real run (simulation). Four terrains; in each row the same policy runs at five risk settings, from risk-averse on the left to risk-seeking (“risk-affine”) on the right. Going down the stairs (top left), the bold robots on the right tumble while the careful ones on the left walk down. Footage: Robotic Systems Lab, ETH Zurich (the paper’s video).

The 3-D scene is a re-enactment: robots, terrain dimensions and the outcome shown for each setting follow the paper’s text; the motion itself is illustrative. The paper reports this experiment qualitatively, without counts.

Scene 7

Risk-aware, without giving up walking skill

In short: set to neutral, the new method learns as well as standard PPO; across the dial it behaves as a risk-aware controller should; and it beats reward-based ways of adding a “risk dial”.

Experiment A: does distributional learning cost performance?

Ten policies per method were trained with simple MLP networks and a risk-neutral DPPO, then evaluated in a deterministic test world: 72 robots per terrain tile, random headings, commanded to walk forward at 1 m/s.

MethodWhat it isReturn after 20k iterations
PPOstandard, not risk-aware27.86
DPPO (this paper, neutral)distributional critic + risk metric27.62
DSACearlier distributional, risk-capable method21.1

What this means: the extra machinery costs essentially nothing (27.62 vs 27.86, without retuning PPO’s hyperparameters), while DSAC failed to learn proper walking.

Paper Fig. 5. Learning curves for the three methods. DPPO and PPO track each other closely; DSAC stays well below.

Experiment B: turning the dial

For each method the authors swept the risk setting from seeking to averse and measured three things: the return (with the original reward), the share of runs that ended early, and the error between commanded and actual speed. A risk-aware policy should trade these off as the dial turns, and all methods did so in the intuitive direction. Risk-seeking tracked the command better but ended more runs early; risk-averse gave up some tracking to keep the robot safe. Return generally rose toward the averse end: in this environment, high-reward moves that often fail didn’t pay off.

Two baselines put the dial into the reward instead, keeping a normal critic:

  • PPO1 scales the speed-tracking reward by β ∈ [0, 2]. At the most averse setting that reward was zero during training, so the robot learned to stop walking.
  • PPO2 treats the reward terms as a small distribution and re-weights them with the Wang metric. It could only be trained for β ∈ [−0.25, 0.25], and its return was much lower, with many early terminations and a large tracking error.

DPPO-Wang earned clearly more return than both. Ablations of DPPO (Quantile-Huber loss instead of energy distance, λ = 0.95 or 1-step targets instead of N-step, CVaR instead of Wang) all earned less; CVaR did worst, perhaps because it ignores a whole part of the distribution, which seemed to make training less stable. Plain PPO, shown as a flat reference line, isn’t risk-aware and sits slightly above all of them in return.

Paper Fig. 6, left. Return across the dial (seeking → averse). DPPO-Wang (solid blue) sits above the ablations and the reward-shaping baselines PPO1 (orange) and PPO2 (green); the black line is plain PPO. One seed per method.
Paper Fig. 6, middle. Share of runs that ended before the episode was over, across the dial.
Paper Fig. 6, right. Speed-tracking error across the dial. Legend for all three panels below.
Scene 8

On the real robot: the dial changes what it dares

In short: on the real ANYmal, the careful setting mostly stayed put in front of a 41 cm step, and the bold setting was the only one that usually made it up a 27 cm step.

The policy was moved to the real robot with the domain randomisation from legged_gym and used ANYmal’s onboard elevation mapping for the height map. The robot was commanded to walk up a step at 1.2 m/s, and each try was counted as a refusal, a failure (an active attempt that didn’t get up) or a success. There were 45 tries in total.

Setting27 cm refusalfailuresuccess41 cm refusalfailuresuccess
Risk-averse25 %50 %25 %85.7 %14.3 %0 %
Risk-neutral12.5 %75 %12.5 %0 %100 %0 %
Risk-seeking0 %37.5 %62.5 %0 %100 %0 %

What this means: the dial carried over from simulation to hardware. Careful mode holds back at a step that is too high, and bold mode commits, and at 27 cm that commitment usually pays off.

The authors note one inconsistency: at 27 cm the risk-averse setting succeeded more often than the neutral one (25 % vs 12.5 %), although overall the neutral setting attempted the step more often. At 41 cm, none of the settings got up; neutral and seeking failed on every try, while averse mostly refused.

Paper Fig. 7. Real ANYmal at the 41 cm step (top) and the 27 cm step (bottom). From left to right: risk-averse stops short, risk-neutral and risk-seeking push into the step; at 27 cm the seeking robot gets up.
Real run. The same step, two dial settings. Top: risk-seeking (“risk-affine”) pushes into the step, gets its front legs up and slips back. Bottom: risk-averse walks up to it and backs away. The video doesn’t say which of the two step heights this is. Footage: Robotic Systems Lab, ETH Zurich (the paper’s video).

Two small discrepancies in the paper, reported as found: the text says 8 tries per setting at 41 cm and 7 at 27 cm, but the percentages (12.5 % = 1 of 8; 85.7 % = 6 of 7) fit 8 tries at 27 cm and 7 at 41 cm. The table labels the averse setting β = −1.5 and the seeking setting β = +1.5, the opposite of the Wang sign convention in the method section (β > 0 averse). We follow the labels “averse / neutral / seeking”.

Wrap-up

Why it matters, and what’s still open

In short: risk becomes something a person or a planner can dial in at run time, without retuning rewards, but choosing the setting and evaluating “risk-awareness” are still open.

What’s new. The locomotion controller models risk explicitly, which deployed controllers didn’t do. Risk-awareness comes from risk-sensitive advantage estimates taken from a learned value distribution, not from reward engineering. And one parameter, changeable during deployment, spans careful to bold.

What it enables. The authors point to teleoperation and navigation. An amateur operator could drive the robot in risk-averse mode, and an expert or a planner could switch to neutral or seeking to use the robot’s full ability when a mission needs it.

What’s still open, in the authors’ words from the conclusion: procedures for evaluating risk sensitivity, and theory. How does the choice of risk metric shape the policy, and how well does the learned value distribution match the true distribution of returns? Integrating DPPO into a navigation system that picks the risk setting itself is named as future work. Also note the scale of the evidence: experiment B used a single seed per method, and the hardware test had 45 tries on two step heights.