CoRL 2024 Workshop · 2024 · an animated walkthrough

Learning a Risk-Averse Locomotion Policy from Uncertainty Estimates

The paper in 106 seconds · narrated · sound on
Transcript

This robot’s front camera is taped over. Told to walk forward, it refuses. On a box whose edges it can’t see, it won’t step off. Once its side cameras have seen the floor, it walks on. How does a robot learn to know what it doesn’t know?

Legged robots like ANYmal explore tunnels and rough terrain, reading the ground as a height map: a grid of ground heights around their feet. But maps have holes. Past a box edge, the ground is hidden. A controller that never learned to doubt guesses “flat”, and walks off. Earlier controllers copied a teacher that saw the true terrain in simulation. It never doubted, so neither did they.

This paper gives the controller a sense of its own doubt: how unsure its guess of the ground is, from noisy input, or from something it has never seen. At a hidden pit, both controllers guess flat ground. Only the new one marks that guess as deeply uncertain.

It learns this in simulation, where the truth is known. Then reinforcement learning, trial and error with rewards, teaches it what to do when unsure. It learns to slow down, take smaller steps, stop, or even back away.

In simulation, imitation alone never refused to enter a noisy zone. After practice it did, and explicit uncertainty made it careful sooner. On the real ANYmal-D, with its camera taped, it refused to walk until its side cameras had seen the floor. Then it went on.

A step toward robots that know the limits of their own senses. Next: turning to look, instead of just stopping. “Learning a Risk-Averse Locomotion Policy from Uncertainty Estimates.”

Robot footage: from the authors’ paper video (Thomas B. Brunner et al., Robotic Systems Lab, ETH Zurich and TU Munich). Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

Researchers at ETH Zurich and TU Munich taught a four-legged robot to notice when it can’t trust what it sees, so it slows down, stops or backs away instead of walking blindly over an edge.

  1. Why this matters

    Four-legged robots like ANYmal hike rough outdoor terrain and explore underground tunnels, places that are hard or dangerous for people. They “see” the ground with depth cameras and turn it into a height map around their feet.

  2. What makes it hard

    That map has holes and errors: the ground behind an edge is hidden, a lens gets covered, reflections add noise. It’s like crossing a dark room. The sensible thing is to slow down where you can’t see, not to stride on as if the floor were there.

  3. What people did before

    Learned controllers that copy a simulated “teacher” who sees the true terrain walk very robustly, but the teacher never doubts, so the student never learns to doubt either. Ways to measure a neural network’s uncertainty existed, mostly for images, and robot-learning work using them mostly stayed on simple simulated tasks.

  4. What this paper does

    It gives the controller a sense of its own doubt. Next to its guess of the terrain, it estimates how unsure that guess is, learned in simulation where the truth is known. Then it practises by trial and error, so it learns what to do when it is unsure.

  5. What they showed

    In simulation, only the controller that practised after imitating ever refused to walk into a zone of scrambled sensor data, and with explicit uncertainty it turned careful at lower noise than without. On the real ANYmal-D it wouldn’t step off a box whose edges it couldn’t see, and with its front camera taped over it waited until its side cameras had seen the floor.

  6. Why it’s a step forward

    It shows uncertainty estimates changing how a real legged robot walks, not just a toy problem. The honest limit: today the robot only stops. The authors’ next step is to let it turn and look before going on.

Words used below
Height map
a grid of ground heights around the robot, built from its depth cameras
Belief
the controller’s internal guess of the terrain, built from noisy inputs
Data uncertainty
doubt because the input itself is bad (noise, holes)
Model uncertainty
doubt because the network has never seen anything like this
Reinforcement learning (RL)
learning by trial and error from a reward
Teacher–student training
a student network copies a teacher that had perfect information
1 / 8
height map / belief uncertainty risk-averse controller (ours)
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough.

Scene 1

Why it matters

In short: legged robots already walk through rough places using what their cameras show them, so the next question is what they should do when the cameras can’t show enough.

Learning-based controllers, trained with reinforcement learning in simulation, have made four-legged robots remarkably capable. The paper builds on ANYmal work from the Robotic Systems Lab that walked over rough terrain, handled harsh conditions and recovered from disturbances, including the DARPA Subterranean Challenge. The same kind of learning has driven progress in robot hands and navigation.

These controllers take two kinds of input. Proprioception is the robot’s sense of its own body: joint positions and velocities, its recent actions, the phase of each leg, and the walking command. Exteroception is what it senses about the outside world: height samples taken from a height map (the paper calls it an elevation map) centred on the robot and built from its depth cameras.

The paper’s starting point is a simple observation: these systems often behave overconfidently, because they have no idea how uncertain their own predictions are. A robot that is about to walk into ground it cannot see should behave differently from one walking on ground it can see clearly.

Real run. ANYmal-D in the lab with the controller from this paper. The inset shows its height map (coloured surface), its height samples (dots) and the operator’s command (red arrow). Footage: Brunner et al., Robotic Systems Lab, ETH Zurich / TU Munich.
Scene 2

Blind spots in the map

In short: the height map is often incomplete or scrambled, and a controller without a sense of doubt fills the gap with a confident guess and keeps walking.

Real sensors fail in ordinary ways. When the robot stands on a box, the ground just below the edge is occluded: no camera ray can reach it, so the map has a gap exactly where the next step would land. A camera can be covered or dirty, so everything in front is unknown. Reflections, dust or vegetation can scramble a patch of the map with noise.

A controller that has no representation of uncertainty has only one option: treat whatever its internal belief says as the truth. In the paper’s pit example (Fig. 4, Scene 5), that belief is simply “flat ground”, with nothing to say how confident it is.

Animated. Rays from the camera stop at the edge, so a strip of ground right below it is never seen. The lower the camera, the wider this blind spot. Geometry illustrative.

The first two cases are exactly the paper’s real-robot tests (Scene 8). The noisy patch stands in for the noise the authors inject in simulation.

Scene 3

Copying a fearless teacher

In short: the best earlier approach trains a student to copy a teacher who could see the true terrain, and it inherits the teacher’s confidence even when its own view is poor.

The paper extends Miki et al. (2022), which uses privileged learning (also called “learning by cheating”). First, a teacher policy is trained with reinforcement learning in simulation using ground-truth information: the true terrain, contact states, friction and external pushes. A real robot can’t measure any of these. Then a student policy, which only gets the observations a real robot has, learns to imitate the teacher’s actions. This is what lets the policy work in the real world.

The catch, in the paper’s words: the student is limited to imitating a teacher that behaves confidently because it always had perfect observations. When the student’s terrain observations are heavily corrupted by noise, it does not change its behaviour, even though it cannot foresee the terrain in front. Ideally it would be more careful.

Uncertainty estimation itself was not new. Kendall and Gal (2017) separated two kinds of uncertainty for computer vision, and methods such as Monte Carlo dropout and deep ensembles were established. Some works had used such estimates to make robots more robust, but, as the paper notes, mostly on simple synthetic tasks, not on real robots.

Not to be confused with “Learning Risk-Aware Quadrupedal Locomotion using Distributional Reinforcement Learning” (Schneider et al., ICRA 2024), also from the Robotic Systems Lab, which has its own explainer. That work makes the controller consider the spread of possible future rewards, with a risk preference the operator can adjust. This paper instead estimates how unsure the controller is about what it perceives, and learns to be cautious when that uncertainty is high.

Scene 4

Two kinds of doubt

In short: the controller learns to say both “my input is bad here” and “I’ve never seen anything like this”, and adds the two into one uncertainty signal.

The student network has three parts (Fig. 1): a belief encoder that turns the inputs into an internal belief state, a belief decoder that rebuilds the observations from that belief, and an action network that decides the joint commands. New here: the decoder also outputs an uncertainty, and the action network receives it next to the proprioception and the belief. The rebuilt observations themselves are only used for training.

Paper Fig. 1. The student policy. The belief decoder’s uncertainty output goes into the action network. The rebuilt observations (dotted arrow) are only used to train it.

Data uncertainty: an error bar per height sample

For every height it rebuilds, the decoder predicts a mean ŷ and a variance σ² (in practice the log-variance, for stable training). In simulation the true terrain y is known, so the network is trained with the loss from Kendall & Gal:

Ldata = 1/N · Σ [ ½ σi−2 (ŷi − yi)² + ½ log σi² ]  // Eq. 1: large errors are forgiven only if σ is large; large σ costs log σ²

No one labels the uncertainty. The loss itself pushes σ toward the size of the actual error: where the input is noisy or missing, errors are big and the best σ is big.

Animated. As the prediction error grows, the lowest point of the Eq. 1 loss moves to a bigger σ. For a single sample, the loss is lowest exactly when σ equals the error. Worked example of Eq. 1.

Model uncertainty: do several guessers agree?

To estimate model uncertainty, the paper combines Monte Carlo dropout (randomly switching off parts of the network at every pass) with a deep ensemble (several separately trained networks). Running M such forward passes gives M guesses. Where they agree, the network knows this kind of input; where they spread out, it doesn’t. The final belief is the mean of the M passes, and the total uncertainty is:

Var(ŷ) ≈ [ 1/M Σ ŷm² − (1/M Σ ŷm)² ] + 1/M Σ σm²  // Eq. 3: spread of the M guesses (model) + average predicted σ² (data)
Animated. Five copies of the network rebuild the same stretch of ground. On familiar input they agree; on unfamiliar input they fan out, and that spread is the model uncertainty. Illustrative.
Paper Fig. 6. Inside the student. Top: the belief encoder (a recurrent GRU memory with an attention gate on the height samples, as in Miki et al. 2022) and the decoder. Bottom: the ensemble with MC dropout produces M belief states, reconstructions and data uncertainties. Their mean becomes the belief; the variance of the reconstructions plus the mean data uncertainty becomes the uncertainty input to the action network.
Scene 5

Doubt, drawn in 3-D

In short: at a pit it cannot see, both controllers believe the ground is flat, but only the new one also says “I’m not sure”.

Fig. 4 shows the robot in simulation walking toward a pit it cannot observe because of occlusion. Without uncertainty estimates, the controller’s belief shows flat terrain, and it has no way to say how confident it is in that. The uncertainty-aware controller’s belief is still flat, but it correctly marks those samples as uncertain: long bars for the front-left leg, which is closest to the pit, and shorter ones for the front-right leg.

This is the key point. The new controller doesn’t magically see the pit. It knows that it doesn’t know, and that knowledge is an input the action network can learn to use.

Paper Fig. 4. Left: the belief of a controller without uncertainty, flat and with no confidence attached. Right: the uncertainty-aware controller. Its belief is also flat, but the stretched samples near the front legs show high uncertainty over the unseen pit.
Scene 6

Training recipe

In short: a harder simulated world, with scrambled sensors and ground you can’t cross, plus a third training step where the student practises on its own.

A harder world. To learn when its data is bad, the controller must experience many kinds of bad data. The simulated environment is divided at random into sub-regions, and each sub-region gets its own randomly sampled noise configuration on the height samples. The terrain also now includes non-traversable obstacles (Fig. 5d), so walking on is not always the right answer.

Animated. Each sub-region distorts the height samples in its own way. The paper doesn’t list the noise types, so the four shown here are examples. Illustrative layout.
Paper Fig. 5. Training terrains with adjustable difficulty. The red dots are the robot’s height samples. The first three can be walked over; the last one (pit and walls) cannot.

Three steps (Fig. 7):

  1. RL of the teacher. Trained with PPO (a standard reinforcement-learning algorithm) on privileged observations.
  2. Supervised learning of the student. A behaviour loss makes it copy the teacher’s actions (behaviour cloning). A reconstruction loss against the privileged ground truth teaches the decoder to rebuild the terrain and, through Eq. 1, its uncertainty.
  3. RL of the student. Starting from those weights, the student is fine-tuned with PPO on noisy observations. It maximises the task reward given the noisy observations, its belief and the uncertainty of that belief (Eq. 4).
maxπ E[ R | π(õt, bt, Var(bt)) ]  // Eq. 4: õ noisy observations, b belief, Var(b) its uncertainty
Paper Fig. 7. The three training steps, all in simulation. The student’s weights from step 2 are the starting point for step 3.
Scene 7

Practice makes careful

In short: copying can’t teach caution the teacher never showed. Trial and error can, once the student can feel its own uncertainty.

Why is step 3 needed? After step 2, the student has an uncertainty signal but has only ever been rewarded for matching a teacher that never hesitated. Nothing yet connects “high uncertainty ahead” to “act differently”. Fine-tuning with reinforcement learning lets the student find its own behaviour, and the results (Scene 8) show this step is necessary: without it, the policy never refused to enter the uncertain region.

The behaviours the authors saw emerge: walking slower, taking smaller steps, avoiding the uncertain region, and even stepping backwards away from it.

Animated. As the amber uncertainty bars grow ahead, the robot slows and shortens its steps, then stops. If it gets too close, it steps back. Illustrative motion of the behaviours the paper reports.
Simulation. Left, “SL only” (imitation, no fine-tuning); right, “SL + RL” (this paper). Both walk among pits. Near an edge, the imitation-only robot steps in (its view goes dark), while the fine-tuned robot holds back; the green bars over the pit are its uncertainty estimates. Footage: Brunner et al., Robotic Systems Lab, ETH Zurich / TU Munich.
Scene 8

Results

In short: only the fine-tuned controller learned to refuse, explicit uncertainty made it refuse sooner, and the real robot refused to walk into areas it couldn’t see.

Simulation

The robot walks over flat ground toward a goal, with a region of uncertainty (noisy height data) in the middle. Each policy is run NT = 20 times at each noise intensity. The authors count the runs NR where the robot refuses to advance into the region and report the relative risk-aversity R = NR / NT. R = 1 means it avoided the region every time.

Animated. How R is counted. The end points are clear in Fig. 2: with no noise, no policy refuses. At the strongest training noise, both fine-tuned policies always refuse, while the imitation-only policy never does.
Paper Fig. 2. Risk-aversity against noise intensity (0 = no noise, 1 = the strongest noise used in training). Left: imitation plus fine-tuning (blue) against imitation only (green, flat at zero). Right: fine-tuned policies with uncertainty (blue) and without it (red, same architecture).
Comparison (Fig. 2)Result as stated in the paper
Imitation only (SL) vs. imitation + RL fine-tuningFine-tuning with RL was necessary: the imitation-only policy never refused at any noise level
Fine-tuned, without vs. with explicit uncertaintyWith uncertainty, the policy avoided the uncertain region at a lower noise intensity

What this means: the sense of doubt alone isn’t enough; the robot needs practice to learn what to do with it. Once it has both, it becomes cautious sooner than a policy that has to infer its doubt implicitly. The paper gives these results as plots only; the curves in Scene 8 are redrawn approximately from them.

Real ANYmal-D

Standing on a box. The box’s edges were hidden from the cameras by occlusion, leaving gaps in the height map. The policy did not step off the box: it treated the area around it as uncertain and refused the operator’s command.

Real run. ANYmal-D on a box. The inset shows the command arrow pointing off the edge; the robot stays on the box. Footage: Brunner et al., Robotic Systems Lab, ETH Zurich / TU Munich.

Front camera taped over. With the front depth camera covered, the area in front of the robot became unknown. The policy refused to walk forward until the area had been seen by the side cameras. Once it had, the robot went on safely.

Real run. Front camera covered: the command arrow points forward, but the robot holds its position. Footage: Brunner et al., Robotic Systems Lab, ETH Zurich / TU Munich.
Paper Fig. 3. Real-robot tests: occlusion at a box edge, the front sensor covered, and after the side camera has observed the area. Insets show the height map, the height samples and the command arrow.
Real testWhat the robot could not seeWhat the controller did
Standing on a boxground just past the edges (occlusion)did not step off
Front depth camera tapedeverything in frontrefused to walk forward until the side cameras saw the area, then proceeded

What this means: the caution learned in simulation carried over to the real robot. It overruled the operator’s command exactly when its map couldn’t support the step. These are qualitative demonstrations; the paper reports no success rates for the hardware tests.

Wrap-up

Limits and next steps

In short: the robot learned to be careful; the next step is to be curious, turning to look instead of just stopping.

  • It stops rather than looks. When the area ahead is uncertain, the controller stops or backs off. The authors name active perception as future work: turning to observe unknown areas before going on.
  • A short workshop paper. Results are given as plots (simulation) and demonstrations (hardware). There are no tables of success rates, and the paper does not list network sizes, noise parameters or reward terms.
  • Caution has a cost. A risk-averse robot may refuse commands. In the real tests that was the desired outcome, but in some missions an operator may need to override it.

The contributions, in the paper’s terms: (1) a way to estimate both model and data uncertainty in a locomotion controller by using ground truth from simulation, and (2) evidence that reinforcement learning can use these estimates to produce risk-averse behaviour, in simulation and on a real ANYmal-D.