IEEE/RSJ IROS · 2023 · an animated walkthrough

MEM: Multi-Modal Elevation Mapping for Robotics and Learning

The paper in 106 seconds · narrated · sound on
Transcript

A robot driving through a vineyard keeps a map of the ground around it, to decide where it can go. But a map of heights only sees shapes. A person lying in tall grass is just another bump, as high as the grass itself.

Cameras and neural networks can tell road from grass, or a person. But a photo has no depth, and every robot carries a different mix of sensors.

This paper, from ETH Zurich, keeps the fast height map, and gives every cell extra layers: colour, what is there, and learned features. Like a spreadsheet of the ground, with one sheet per kind of information.

3-D points, from laser scanners and depth cameras, simply drop into the cell below them. Plain photos work the other way round. Each map cell is traced into the picture, and a line-of-sight check stops cells hidden behind the box from stealing its label. A fusion rule, how new readings mix with old ones, keeps the map steady when the camera network flickers. Small plugins then read the layers, for example to trace vineyard tree rows with a learned network.

It all runs on the graphics chip. A full update takes 23.6 milliseconds on the robot’s small onboard computer, about 42 times a second, and the extra layers add just 1.5 of those milliseconds. And the person in the grass? The class layer marks them clearly, where height alone could not.

One open-source map, for legged and wheeled robots, that holds geometry and meaning together, and is ready to feed learning. Multi-Modal Elevation Mapping for Robotics and Learning.

No result footage is used: the video is built from the page’s animations and the paper’s figures (CC BY 4.0). Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video. Animations are schematic.

The story in plain words

A team from ETH Zurich and Oxford taught a robot’s map of the ground to remember not just how high the ground is, but what it looks like and what it is, fast enough to run on the robot’s own computer.

  1. Why this matters

    Robots that work outdoors, like a four-legged robot on a city pavement or a small farm robot in a vineyard, keep a map of the ground around them to decide where to step or drive. Knowing what the ground is helps: a robot that can tell road from grass and mud can choose the easier path.

  2. What makes it hard

    A map of heights sees shapes, not meaning, like feeling the ground with a walking stick. A person lying in tall grass is just another bump. Meaning comes from cameras, but a picture has no depth, and every robot carries a different mix of sensors with different noise and speed.

  3. What people did before

    Height maps (one height per grid cell) are the usual choice for ground robots, and the team’s own graphics-chip version was already fast, but it stored geometry only. Maps that also stored meaning existed, but they were fixed to class labels, ran on the main processor, needed depth-aligned images, or were not public.

  4. What this paper does

    It gives every map cell extra layers, like a spreadsheet of the ground with one sheet for height, one for colour, one for “what is here”, one for learned features. Data comes in as 3-D points or as plain photos, a menu of rules decides how new readings mix with old ones, and small plugins turn layers into answers. They call it MEM, the Multi-Modal Elevation Map.

  5. What they showed

    A full update takes 2.6 ms on a desktop graphics card and 23.6 ms on a Jetson Orin, the small computer inside the robot: about 42 times a second. The extra layers add 1.5 ms of that. On robots it coloured a map from three cameras, found a person lying in tall grass, and traced vineyard tree rows.

  6. Why it’s a step forward

    One open-source tool now serves legged and wheeled robots and many tasks, and its layers can feed learning methods directly. The demos are shown as pictures, not scored: the paper measures speed and memory, not how accurate the labels are.

Words used below
elevation map
a grid seen from above; each cell stores one ground height (“2.5-D”)
layer
one more value per cell: colour, class probability, or a feature
semantic segmentation
a network labels every pixel with what it shows
feature
a list of numbers a network uses to describe an image patch
fusion
combining a new reading of a cell with what the map already held
GPU
a graphics chip that does thousands of small calculations at once
1 / 7
geometry, point clouds camera images multi-modal map (MEM)
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper (CC BY 4.0); the text is a plain-language walkthrough.

Scene 1

Why a map needs more than height

In short: robots that work outside need to know not only the shape of the ground but also what it is, and height maps alone cannot tell them.

Ground robots such as the four-legged ANYmal or wheeled farm robots carry laser scanners, stereo cameras or depth cameras. Each sensor gives a different kind of data with its own noise, which is why robots merge them into one map of their surroundings. For robots that move on the ground, the favourite is an elevation map: the ground seen from above, cut into square cells, with one height stored per cell.

Heights are enough to decide whether a step is too tall. They are not enough to tell concrete from grass or mud, which matters when a robot should prefer the road, or to notice a person lying in tall grass, whose body is about as high as the grass around it. The authors’ goal was to let the same map also hold colour, class labels, learned image features and anything else a task needs, without losing the speed that made their earlier map useful on real robots.

Paper Fig. 1. The same mapping software on three jobs. Left: a pavement in Zurich, coloured by compressed image features. Middle: a wheeled robot in a vineyard, coloured by class. Right: a garden, coloured by class.
Scene 2

One map, many layers

In short: the paper keeps the fast height map ground robots already use and adds as many extra layers per cell as the user wants.

Robot maps come in three kinds. 2-D occupancy grids only say whether a cell is free or blocked, which is cheap but useless on steps and slopes. 3-D voxel maps store every block of space, so they can represent a table or an overhang, but they cost much more memory and computing time. 2.5-D elevation maps store one height per cell: they handle stairs, steep slopes and rubble, and are cheap enough to update many times a second, which is why they are popular for legged and wheeled robots.

Animated. The same slice of ground stored three ways. Watch the 2.5-D map fill the space under the table: it keeps one height per cell, which is the price of being fast.

The team’s earlier GPU elevation map (Miki et al., 2022) was already fast but stored geometry only. MEM keeps that grid and adds multi-modal layers. The layers don’t care what they hold: a layer can be red, green or blue colour, the probability of “grass”, or one of 90 feature numbers from a vision network. The map is centred on the robot and moves with it.

FeatureMEM (this paper)Elmap [Miki 2022]Selmap [Ewen 2022]Multitask [Gan 2021]Maturana 2018
Class probabilitiesyesnoyesyesyes
Learned featuresyesnonoyesno
Real timeyesyesyes(yes)yes
Plain (monocular) imagesyesnononoyes
ROSyesyesyesyesyes
GPUyesyesnono?
Open sourceyesyesyesyesno

What this means: each earlier system has some of these properties; MEM is the only one in the paper’s comparison (Table I) that has all of them.

The paper is fair about what the others do well. Maturana et al. fuse class probabilities from images into a robot-centred height map with the same kind of Bayesian update MEM uses, but their code is closed. Gan et al. build a large 3-D semantic map, but on the main processor, and adding a new kind of information means retraining their whole multi-task network. Ewen et al. build a semantic height map from one RGB-D camera, focused on class probabilities.

Paper Fig. 2. The whole system. Images (purple) and point clouds (orange) are matched to cells, fused into layers, and then read by plugins such as line detection or traversability estimation. Everything runs on the GPU.
Scene 3

Points in

In short: a 3-D point goes to the cell underneath it, and anything the point carries (colour, class, features) goes into that cell’s layers.

A point cloud is a set of 3-D dots measured by a laser scanner or depth camera. MEM accepts point clouds with any number of extra fields per point, for example a colourised cloud or a semantic point cloud made by combining a depth image with a network’s per-pixel class probabilities. Each point is matched to a cell by its horizontal position, all points in parallel. Its height updates the elevation layer exactly as in the earlier map; its extra fields update the other layers.

When several points fall into the same cell in one scan, their values are first averaged, layer by layer:

at,j = (1 / Nj) · Σi mi · [mi in cell j]  // mi: the i-th measurement, Nj: points in cell j
Animated. Points land in the cell below them; the highlighted cell keeps the average of whatever landed in it this scan.

Point clouds are big. The paper lists what its sensors deliver per message:

SensorPoints per cloudRate
ZED 2i stereo camera230,4003 Hz
Velodyne VLP-16 lidar300,0001 Hz
Depth cameras407,04015 Hz

What this means: hundreds of thousands of points arrive several times a second, which is why the matching is done for all points at once on the GPU.

Scene 4

Images in

In short: a plain photo has no depth, so MEM asks the opposite question: for each map cell, which pixel shows it, and can the camera really see it?

A camera image can carry colour, a class per pixel, or a feature vector per pixel, but it does not say how far away each pixel is. MEM uses the heights already in the map instead. For every cell inside the camera’s field of view, it computes where the cell lands in the image with the standard pin-hole camera model (using the camera’s known lens parameters and position on the robot) and copies that pixel’s value. Going from cells to pixels is the cheaper direction, because a robot-centred map usually has fewer visible cells than the image has pixels.

A cell hidden behind a box or a wall must not receive the box’s label. So for each cell MEM walks the cells between the camera and that cell, picked with Bresenham’s line algorithm (a classic way to list the grid cells a straight line passes through). The cell counts as visible only if every one of those cells is lower than the line of sight at that point. Every cell is checked in parallel.

Animated. The cells between the camera and the target light up one by one. When the line passes the tall box, the target is marked hidden and keeps its old value.
Paper Fig. 3. The idea from the side. Each cell in view is linked to a pixel of the labelled image; the ray to a cell behind the box hits the box first (orange circle), so that cell is skipped.

Demo: colouring the map from three cameras

Because images no longer need their own depth, a robot can colour its map from ordinary colour cameras while its depth comes from other sensors. Earlier approaches needed a colourised point cloud or a colour image aligned with a depth image. On the ANYmal C legged robot, the paper used three colour cameras (Sevensense Alphasense Core, 1080 × 1440 pixels at 9 Hz) for colour and four Intel RealSense D435 depth cameras (480 × 848 at 15 Hz) for height.

Paper Fig. 7. Three camera views (top) painted onto the height map (bottom). The curb’s yellow-black stripes land on the right cells.
Scene 5

Fusion rules

In short: readings are noisy and the world changes, so each layer picks one of four rules for mixing a new reading with the value it already holds.

Once a reading is matched to a cell, it has to be combined with the cell’s previous value. Different kinds of data need different rules, so the user picks one per input source and per layer, and can send several sources into the same layer or keep them separate. New rules can be added.

  • Latest. Keep only this scan’s value (the average from Section 3). Works for any data, but ignores the past and so is the noisiest.
  • Exponential averaging. Blend the new value with the old one using a fixed weight w set by the user. Robust to noise and usable for any numbers, but the result is not a probability.
  • Gaussian Bayesian inference, for feature vectors: each feature number is treated as a noisy measurement of a hidden mean; the update has a closed form.
  • Dirichlet Bayesian inference, for class probabilities: keep a running count per class and read probabilities off the counts.
θt = w · at + (1 − w) · θt−1  // exponential averaging
Animated. The grey line (Latest) jumps with every reading; the exponential average smooths the noise and still catches up when the true value changes.

For features, the mean after N readings is a weighted mix of the prior mean and the average of the readings, and the spread shrinks as readings accumulate:

μN = σf² / (Nσ0² + σf²) · μ0 + Nσ0² / (Nσ0² + σf²) · μML    σN² = σf²σ0² / (Nσ0² + σf²)  // μ₀, σ₀²: prior; σ_f²: known measurement variance; μ_ML: mean of the readings
Animated. Each new feature reading pulls the belief toward the data and makes it narrower, computed with the paper’s formulas.

For classes, the cell keeps one count αk per class. Each reading adds its class probabilities to the counts, and the cell’s probability for class k is its count divided by the sum of all counts:

αt = αt−1 + Σi mi    p(class k) = αk / Σ α  // a Dirichlet prior with a categorical likelihood

Because the counts remember every past reading, this rule changes its mind gradually, while exponential averaging forgets old readings quickly. Paper Fig. 4 shows this on a cell that first says “class 0” and then keeps reading “class 2”. To save memory with many classes, the map can store only a chosen subset of classes or the top k.

Paper Fig. 4. Top: the readings. Middle: exponential averaging hands the cell over to class 2 within a few steps. Bottom: the Bayesian rule still favours class 0 at t4, because it remembers the earlier readings.

Demo: a person in tall grass

In a vineyard, a stereo camera (360 × 640 pixels) fed images to Detectron2, a segmentation network that outputs class probabilities for every pixel. Together with the depth image these became a semantic point cloud, fused with the Dirichlet rule. The person lying in the grass cannot be told apart by height, but stands out in the class layer, so a local planner could treat it as an obstacle.

Paper Fig. 8. The person (orange) is about as high as the vegetation (red) in the height map; only the class layer makes them visible.
Scene 6

Plugins turn layers into answers

In short: small add-ons read any mix of layers and either write a new layer or hand a result to other software, which makes the map a ready-made input for learning.

Post-processing plugins run when an event triggers them, at a rate the user sets, and are configured with a few lines of Python. The paper extends the plugin system of the earlier map so plugins can also read colour, class and feature layers. Since the layers already sit on the GPU, a plugin does not have to copy data around. A plugin can create a new layer, change an existing one, or send a result elsewhere, such as the tree lines below.

Demo: finding vineyard tree lines

A wheeled robot with one RGB-D camera (360 × 640 pixels) drove through special crops. Features from a self-supervised vision transformer (DINO) were fused with the Gaussian rule. A network then read the height and feature layers: an ERFNet (a compact image-segmentation network) produces one weight map per tree line, and a least-squares layer fits each line as a second-order polynomial in the robot’s frame, following Van Gansbeke et al. (2019). The authors note that such lines can guide the robot even when its own position estimate is inaccurate. A second plugin compresses the feature layers into a three-channel PCA layer that people can look at.

Animated. As more weighted samples arrive, the least-squares fit settles on a smooth second-order curve.
Paper Fig. 5. The vineyard pipeline. Ninety feature layers sit next to height and colour; the line-detection plugin outputs the left (red) and right (blue) tree lines, and the PCA plugin writes a new layer back into the map.

The paper gives no accuracy numbers for the tree lines; the demo shows that the pipeline works end to end. In concurrent work, MEM was also used to fuse visual traversability estimates (Frey et al., 2023).

Scene 7

Fast and small enough for the robot

In short: the extra layers barely slow the map down; it updates about 42 times a second on the small computer inside the robot.

The map is written in Python with CuPy, which lets the team write custom CUDA programs for the GPU; ROS carries the messages, and a C++ wrapper speeds up converting them. The test used a 10 m × 10 m map with 4 cm cells (250 × 250), one ZED 2i camera at 3 Hz (360 × 640 pixels) and a small segmentation network (LR-ASPP with a MobileNetV3 backbone), timed over 300 updates.

Step (point cloud of 230,400 points)RTX 4090 [ms]Jetson Orin [ms]
Point transform & height-error count0.047 ± 0.0130.695 ± 0.246
Drift compensation0.108 ± 0.0930.915 ± 0.551
Height update & ray cast1.341 ± 0.10717.524 ± 2.417
Multi-modal update0.306 ± 0.1151.536 ± 0.508
Overlap clearance0.226 ± 0.1000.970 ± 0.186
Traversability0.538 ± 0.1591.936 ± 0.200
Normal calculation0.030 ± 0.0300.060 ± 0.112
Total2.596 ± 0.26023.635 ± 2.754

What this means: 2.6 ms and 23.6 ms are 385 and 42 updates per second. The new multi-modal step is about 12 % of the time on the desktop GPU and 6.5 % on the Orin; merging the height points is still the slow part.

Compared with Ewen et al., MEM runs at about the same speed with four times as many cells. Adding more layers costs time in a straight line: with 20 layers the whole update on the Orin takes 28.9 ms with exponential averaging and 29.0 ms with the Bayesian rule.

Paper Fig. 6. Time for the fusion step on a Jetson Orin as layers are added. It grows steadily, and the two rules cost about the same at 20 layers.

Outside the map itself, the slowest part of one setup was the segmentation network: LR-ASPP took 31 ms per image. MEM does not depend on a particular network, so users can trade speed for quality.

What is storedMemory
Point cloud, 100,000 points with x, y, z, RGB and 5 class channels7.2 MB
Earlier elevation map (Miki 2022), 200 × 200 cells2.2 MB
MEM, 200 × 200 cells with the same channels3.8 MB

What this means: the extra layers cost 1.6 MB, less than a quarter of one 100,000-point cloud with the same channels, and keeping only the top-k classes would shrink it further.

Wrap-up

Limits and what’s next

In short: MEM is a general, open tool, and the paper proves speed and memory; how good the labels are depends on the networks feeding it.

  • Shown, not scored. Colouring, the person in the grass and the tree lines are shown as pictures. The paper reports no accuracy numbers for them; what it measures is speed and memory.
  • Images need a height map first. Photos carry no depth, so image layers depend on heights already built from range sensors, and a wrong height can send a pixel to the wrong cell.
  • Still 2.5-D. One height per cell cannot store overhangs such as a table top or branches above a path.
  • Only as good as its inputs. Classes and features come from the upstream networks; the map fuses them but cannot fix what they get wrong.

The authors’ conclusion: MEM gives ground robots one easy, efficient map that holds geometry together with colour, classes and features, and they released it as open source (elevation_mapping_cupy) for the research community.