IEEE Robotics and Automation Letters · 2022 · an animated walkthrough

Reconstructing occluded elevation information in terrain maps with self-supervised learning

The paper in 115 seconds · narrated · sound on
Transcript

A four-legged robot maps a mine as it walks. Behind rocks and edges its map has holes; here a network fills them in, in red. With the holes filled, the robot can plan further ahead. But how do you learn to fill holes when no real map is ever complete?

Robots plan their route on an elevation map: a grid seen from above, where each cell stores the height of the ground. The sensor sees along straight lines, so behind rocks the map stays blank, and a careful planner stops at the first hole.

Classic fixes draw a smooth line across each gap. On stairs, that turns the steps into a ramp. A person would guess better: ground hidden behind a crest is probably lower, or the sensor would have seen it. A network could learn such clues, but it needs complete maps to learn from, and real robots never record one.

The trick: hide even more of a real map, as seen from a random viewpoint, and grade the network only where the height was measured. That is self-supervised learning: the data makes its own quiz. The imaginary sensor is raised until at most half the map is hidden. A standard image-repair network fills every gap; measured cells are kept, only the holes use its guesses.

On four real datasets, from a Swiss mine to a lunar-like volcano field, the error in the filled cells fell by 52 to 82 percent. Here it is on the robot’s recorded mine data: blue was measured, red is the network’s fill. Where complete answers do exist, learning from practice holes came close to learning from the answers. It runs 30 times a second on a GPU, and in simulation the robot planned further ahead and stopped less (illustrated).

No perfect maps needed: a robot could keep learning from the maps it makes while it explores.

Result footage: Maximilian Stölzle (first author), from the two videos linked in the paper — inference in the Gonzen mine and motion planning with ANYmal. Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

A walking robot’s map of the ground always has blind spots behind rocks and step edges; this paper teaches a neural network to fill them in using only the robot’s own incomplete maps, by hiding extra patches and practising to guess them back.

  1. Why this matters

    Legged robots like ANYmal and planetary rovers choose their route on a map of the ground seen from above, where every square stores how high the ground is. A safe, efficient route needs that map to be complete.

  2. What makes it hard

    Sensors see along straight lines. Behind a rock, past a drop-off or over a crest, the ground is in shadow, like the ground behind a wall you are looking over. Reflections, dust and blank surfaces punch more holes into the map.

  3. What people did before

    Planners either treat holes as no-go zones, so they can only plan up to the next hole, or fill them with smooth lines and image-repair tricks. Those turn stairs into ramps. Learning a better fill needs complete “answer” maps, and real robots never record one.

  4. What this paper does

    It turns the robot’s own holey maps into practice questions. From a random imaginary viewpoint it hides every cell that viewpoint could not see, then trains a network to fill these new holes, scoring it only where the true height is known. It works like a fill-in-the-blanks test made by covering words in a text you already have.

  5. What they showed

    On four real datasets (stairs, an obstacle course, a mine and a lunar-like volcano field) the error in the filled-in cells dropped by 52–82 % compared with the best classic method. The network handles a 12 m × 12 m map 30 times a second on a GPU, and in simulation the robot could plan further ahead and stop less.

  6. Why it's a step forward

    No hand-made answer maps are needed, so a robot could keep learning from the maps it records while it explores. Honest limits: it also learns to copy sensor noise, it does not yet say how sure it is about each cell, and the planning test was an early study in simulation.

Words used below
Elevation map
a top-down grid; each cell stores the ground height (the paper’s “DEM”).
Occlusion
ground the sensor cannot see because something is in the way.
Inpainting
filling in missing parts of an image, here of a height map.
Ray casting
drawing straight lines from a viewpoint to test what it can see.
Self-supervised learning
training where the data provides its own questions and answers.
U-Net
an image network that shrinks the map to see context, then grows it back.
1 / 7
measured by the sensor practice holes (ray casting) network’s fill (ours)
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper (CC BY 4.0); the text is a plain-language walkthrough.

Scene 1

Every terrain map has blind spots

In short: robots plan on a height map of the ground, and that map is always full of holes wherever the sensor could not see.

Mobile robots that go somewhere on their own need a map of the ground around them to judge where they can walk or drive and to plan a path. A common and compact choice is a 2.5-D elevation map (the paper says DEM, digital elevation map): a regular grid seen from above in which each cell stores one number, the height of the ground there. Robots build it from depth sensors such as a laser scanner (LiDAR) or stereo cameras, and the robot-centric version used on ANYmal keeps the robot at the centre of the grid as it moves.

For safe and efficient planning you want this map complete and accurate. In practice it never is. Rocks, obstacles and convex terrain such as a crest or a step edge hide patches of ground from the sensor’s viewpoint. On top of that, depth measurements fail on reflections, in dust, when stereo matching fails or on textureless surfaces, which leaves further gaps.

Animated. Why holes appear, and why they are not pure mystery: anything hidden behind the crest must lie below the grazing line of sight, or the sensor would have seen it. Schematic terrain.

The authors’ observation: people look at such a map and guess what hides in the holes quite well. We know what stairs, rubble and hills usually look like, and we intuitively use the line of sight. Classic hole-filling methods use neither. The goal of the paper is a method that does, and that can learn from the data robots actually collect.

Paper Fig. 2 (left panel). ANYmal in the Gonzen mine. The measured map is shaded white-to-blue by height; the network’s fill of the missing parts is overlaid in white-to-red. For clarity, only map tiles with less than 85 % missing cells were filled.
Scene 2

Smooth guesses break on real terrain

In short: existing fixes either avoid the holes or smear smooth surfaces across them, and learning a better fill needs complete maps that real robots never record.

What do robots do with holes today? The most common answer in navigation work is to avoid them: the planner simply does not plan through cells without height information. That is safe but short-sighted (scene 7). The other answer is to fill them with methods borrowed from image repair:

  • Interpolation (linear or cubic) draws a smooth surface between the known edges of a hole.
  • Diffusion-based inpainting such as Navier–Stokes and Telea propagates the surrounding values into the hole, following lines of equal value.
  • Patch-based methods copy the best-matching patch from elsewhere. For terrain, Kolter et al. (2009) searched an offline library of terrain patches and also enforced line-of-sight constraints.

These hand-tuned rules do fine on smooth ground but know nothing about what terrain usually looks like. On stairs and boxes they blur edges into ramps. In image inpainting, learned networks (usually U-Nets, sometimes with adversarial training) have long been the state of the art, and Qiu et al. used a GAN to fill voids in large geospatial elevation data.

Animated. Walking down stairs, the sensor cannot see the near part of each tread. A straight line across the gaps makes a ramp; a fill that knows about stairs keeps the steps. Schematic.

The catch for robot maps is training data. A network learns to fill holes by seeing examples with the right answer, which means complete elevation maps. Those exist in simulation, but real maps always have holes. The closest earlier work, Šalanský et al. (2021), built “ground-truth” maps offline with SLAM, which also uses measurements from later in the run and so fills more of the map. The authors point out that such a target inherits pose-estimation errors and flawed depth measurements. The question this paper answers: can we train on the incomplete real maps directly?

Scene 3

No answer key? Make practice holes

In short: take a real map with holes, hide some more of it, and train the network to restore what you hid; the recorded map is the answer key.

The key idea is a new way of doing self-supervised learning on terrain maps. Instead of asking for a complete map, the method treats the partly occluded map the robot recorded as the target. It then makes the input harder by adding artificial occlusion: extra holes that look like the ones a sensor would produce (scene 4 shows how). The network sees the doubly-holed map and must fill it in.

Grading is the clever part. The network’s output is compared with the target only in the artificially hidden cells, where the robot really did measure the height, plus (with less weight) the cells that were visible anyway. Cells that were already missing in the recording are ignored in the loss, because nobody knows their true height. At test time the network fills all holes, including the original ones.

Animated. Which cells count when the network is graded. Orange practice holes are scored against the recorded height; grey holes have no answer and are skipped. Toy map.
Paper Fig. 1. The whole method on one line: maps recorded by ANYmal or the planetary rover become the target; ray casting from a random viewpoint makes the occluded input and its mask; the network’s reconstruction is merged back into the holes; the loss compares against the recorded map.

Because no clean answers are needed, the paper notes that this points towards robots that keep improving from the data they gather while exploring, in the spirit of lifelong learning.

Scene 4

Practice holes that look like real ones

In short: the practice holes come from an imaginary sensor placed at a random spot, and its height is tuned until a sensible share of the map is hidden.

For the trick to work, the practice holes must resemble real occlusion; otherwise the network learns to fill the wrong kind of hole. The authors first tried growing the existing holes (dilation) and deleting random pixels, but neither produced realistic or varied masks. What does: ray casting. They wrote a small, fast C++ component that takes an elevation map and a vantage point and checks, for every cell, whether the straight line from the vantage point to that cell is blocked by higher ground. Blocked cells become holes.

The vantage point is drawn uniformly at random on the map. Its height above the ground, the offset o, decides how much gets hidden: low viewpoints cast long shadows, high ones almost none. An iterative search keeps the occlusion ratio (hidden cells over all cells) between 0.1 % and 50 %:

  1. Sample o uniformly from the range [o_min, o_max] and ray cast.
  2. Too much hidden (> 50 %)? Raise o_min to this o: only higher viewpoints from now on. Too little (< 0.1 %)? Lower o_max to o.
  3. If the range shrinks below 5 cm, widen it again by 5 cm. Stop when the share fits, or after 15 tries.
Animated. The height search, seen from above. Each try recasts the shadows; the gauge shows the occlusion ratio (grey: holes already in the recording, orange: new) against the 0.1–50 % window. Terrain and the starting height range (0–1 m) are illustrative.

The same ray caster also builds the synthetic datasets: it takes a complete generated terrain, places a “robot” at a random position (within ±1.25 m of the centre, sensor 0.2–0.5 m up on hills and 0.2–0.3 m on stairs and boxes) and produces the occluded map and its mask. There, the full answer is known, so supervised and self-supervised training can be compared fairly (scene 6).

Scene 5

A U-Net fills the gaps

In short: a small, standard image-repair network takes the holey map plus a “what is missing” mask and outputs a complete map; only its guesses inside the holes are used.

The network is a U-Net, the design most learned image-inpainting methods build on. It has two input channels: the occluded elevation map and a binary occlusion mask (1 = missing, 0 = known). Missing cells are stored as “not a number”, which a network cannot take, so they are replaced by 0.0 on real data (and by linear interpolation on synthetic data, where that worked better in a selection study). Before the network, the mean height of the known cells is subtracted, and it is added back to the output, so the network only has to reason about shape, not absolute height.

Compared with the original U-Net, the authors kept it light: the inputs are only 64 × 64 cells (not 572 × 572 pixels), so they use 3 down-sampling steps instead of 5 and 64, 128 and 256 feature channels in the hidden layers. The shrinking half sees the wider shape of the terrain; the growing half rebuilds full resolution; skip connections carry fine detail straight across.

L = 10 · MSE(hidden cells) + 1 · MSE(visible cells) + 0.1 · TV(fill)  // MSE = mean squared height error; TV = total variation, a penalty on jagged fills

Training used PyTorch and the Adam optimiser (learning rate 0.0001, weight decay 0.001), for up to 100 epochs or until the validation error on hidden cells had not improved for 50 epochs. At use time the output is composed with the input: every measured cell is kept as it was, and the network’s values are used only where the map had holes.

Animated. Composition: the network predicts every cell, but only its values inside the holes are used. Toy map.
Scene 6

Results

In short: on real robot maps the error in the filled cells drops to between about half and a fifth of the best classic method, and learning from practice holes is nearly as good as learning from complete answers.

The data

  • Synthetic terrain: hills (Perlin noise), regular stairs and randomly stacked boxes; 64 × 64 cells at 4 cm (2.56 m square); 25,000 training and 2,500 each for validation and test.
  • ANYmal with a dome LiDAR: maps of 300 × 300 cells at 4 cm, cut into 16 tiles of 75 × 75. ETH Zurich stairs (26,233 tiles), an ETH obstacle course (37,274) and the Gonzen mine in Switzerland (16,459), recorded during subterranean exploration; split 80/10/10.
  • Tenerife lunar analogue: the ESA HDPR rover testbed with stereo cameras at Minas de San José on the Teide volcano, June 2017. Maps built with GA SLAM using GNSS poses, 600 × 600 cells cut into 64 tiles; 42,600 / 1,000 / 7,950 samples, the test set from a separate side track never seen in training.

Real-world maps

Real maps have no complete answer, so the test is the training trick itself: add ray-cast holes to unseen test maps and measure the error on those cells. Five random seeds; heights in metres.

DatasetBest classicMSE classic (m²)MSE ours (m²)PSNR ours (dB)Change
ETH stairsTelea0.049800.018 ± 0.00130.9 (vs 26.45)−64 %
ETH obstacle courselinear0.346460.166 ± 0.00326.97 (vs 23.77)−52 %
Gonzen minelinear0.144730.026 ± 0.00122.3 (vs 14.87)−82 %
Tenerife lunar analoguelinear0.025770.0110 ± 0.000244.51 (vs 40.81)−57 %

What this means: on every real dataset the learned fill beats all four classic methods (linear, cubic, Telea, Navier–Stokes), cutting the squared error by 52–82 %; for 75 % of test tiles the average error was under 5 cm (stairs), 4 cm (obstacles, mine) or 3 cm (Tenerife).

Synthetic maps: practice holes vs full answers

TerrainBest classicSupervisedSelf-supervisedMSE change (self-sup.)
Hills38.54 dB (cubic)40.0 ± 0.2 dB39.7 ± 0.2 dB−24 %
Stairs32.43 dB (linear)44 ± 1 dB42.2 ± 0.8 dB−90 %
Random boxes9.35 dB (linear)19.0 ± 0.1 dB17.7 ± 0.5 dB−85 %

What this means: PSNR in hidden cells, higher is better (three seeds). Self-supervision is only 0.3–2.3 dB behind training on complete answers, and the gain over classic fills is largest on structured terrain such as stairs and boxes; on smooth hills cubic interpolation is already close.

Paper Fig. 3. Real test tiles: recorded map, map with added holes, Telea’s fill, the learned fill and both error maps. On the stairs, Telea blurs the steps and invents two low dips; in the mine the learned fill recovers the central obstacle better; in Tenerife it correctly guesses that a fully hidden corner lies lower than its surroundings, the line-of-sight clue.
Real data. The trained network running on ANYmal's recorded Gonzen mine maps: measured ground is shaded blue, the cells it fills in behind the walls are shaded white-to-red. Footage: Maximilian Stölzle et al., Robotic Systems Lab, ETH Zurich (video linked in the paper).

The authors also note a side-effect of learning from real data: because the training targets contain sensor noise, the network learns to reproduce similar noise inside the holes.

Speed. For a full 300 × 300 ANYmal map (16 tiles, each resized to 64 × 64, in two batches of 8), the U-Net runs at 3.5 Hz on one core of a laptop CPU and 30 Hz on an NVIDIA TITAN Xp GPU. Telea and Navier–Stokes run at about 50 Hz but are far less accurate.

Scene 7

Planning further ahead

In short: in a first simulation test, a planner using the filled-in map could plan further ahead, so the robot stopped less often to wait for a new plan.

As a preliminary study, the authors fed completed maps into motion planning for ANYmal in simulation. The network was trained on the real ETH stairs data, and the paper highlights how well it carried over to the simulated world. The locomotion and navigation software is the same that team CERBERUS used in the DARPA Subterranean Challenge.

The baseline plans on the raw, incomplete map and never plans through cells without height information, which is common in the literature. It cannot plan far ahead and its paths are not optimised over the whole map, so the robot often has to stop and wait for re-planning. With the reconstructed map, paths reach further and only need small corrections when the real measurements of the hidden area arrive.

Paper Fig. 4 (top row, left). Planning on the incomplete map: the planned path covers only a short stretch in front of the robot.
Paper Fig. 4 (bottom row, left). Planning on the reconstructed map: the planned path reaches much further ahead. The blue line points from the start (right) to the goal (left).
Simulation run. The same stairs walk planned twice: on the incomplete map (top) and on the reconstructed map (bottom, filled cells in red). Footage: Maximilian Stölzle et al., Robotic Systems Lab, ETH Zurich (side-by-side video linked in the paper).

The lanes in scene 7 are an illustrative animation of this behaviour; the paper reports it qualitatively (Fig. 4 and a side-by-side video), without stop counts.

Wrap-up

Limits and what comes next

In short: the method removes the need for complete training maps; knowing how sure it is about each filled cell is the next step.

  • Evaluated on artificial holes. On real data, accuracy can only be measured on cells that were hidden on purpose, because the true height under the original holes is unknown.
  • Learns the noise too. Training targets are real, noisy maps, so the fill reproduces similar noise.
  • No confidence yet. The paper names uncertainty estimation as crucial future work: a planner should trust a filled cell less than a measured one, per cell.
  • Planning is a proof of concept in simulation, not a field test.
  • Ideas from the authors: learn realistic occlusion patterns with a GAN instead of ray casting, and add a discriminator in place of the perceptual losses that need complete answers.