IEEE Robotics and Automation Letters · 2024 · an animated walkthrough

Identifying Terrain Physical Parameters from Vision

The paper in 93 seconds · narrated · sound on
Transcript

Wet floors, soft ground, fresh snow. To a robot’s camera, they all look harmless. Can a robot learn to see how the ground will feel?

Legged robots now work in rescue and inspection. But they only notice slippery or soft ground once a foot slips or sinks. Simulators can set the grip and softness of any patch, but can’t make grass feel real. This paper links the two: it reads the simulator’s own numbers off camera images.

Step one, in simulation: the robot walks over random patches whose friction (grip) and stiffness (hardness) are known. A network with memory learns to read both from one second of motion, separately for each foot.

Walking onto ice, the front foot’s estimate drops first, the hind foot’s a second later. On a real wet whiteboard, the hind foot’s error fell from 0.33 to 0.02. Replaying those steps in simulation confirmed it: the board behaves like friction 0 to 0.2.

Step two, on the real robot: each footstep’s values are painted into earlier camera images, and a small image network learns from them as the robot walks. It also flags pixels it can’t trust. From just 22 images, it found the floor with 82% overlap.

A step toward robots that see slippery and soft ground coming.

Robot footage: Robotic Systems Lab, ETH Zurich, from the paper's supplementary video. Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

A walking robot learns to tell how slippery or soft the ground ahead is just by looking, by first learning in simulation what slippery and soft feel like, then letting its own footsteps teach its camera.

  1. Why this matters

    Four-legged robots now work in rescue sites, mock planetary terrain and industrial plants. What trips them is often not a rock or a step but ground that looks harmless: a wet floor, ice, soft snow, mud. The shape of the ground is fine; how it behaves underfoot is not.

  2. What makes it hard

    A robot finds out that the ground is slippery or soft only once a foot is on it, like walking into a dark room and feeling for ice with your shoe. Seeing it earlier means linking how ground looks to how it feels. Simulators can make ground feel however you like, but can't make grass feel exactly like real grass. Real images have the looks but nobody has written the physics on them.

  3. What people did before

    Robots already estimate grip and softness by touch: with extra force sensors or whiskers, with special probing moves, or (in the lab's own 2022 walking controller) with a memory network that sensed friction as the robot walked. Vision methods learned from footsteps too, but they learned scores such as “traversability” or sounds. A simulator can't reproduce those scores, so policies trained in simulation can't use them.

  4. What this paper does

    It predicts the simulator's own numbers, friction and stiffness, from camera images. Step 1: in simulation, a network learns to read these numbers for each foot from how the robot moves. Step 2: on the real robot, that network labels every footstep, the labels are painted into camera images taken earlier, and a small image network learns while the robot walks. It's like a child who slips on ice once and from then on recognizes ice by sight.

  5. What they showed

    In simulation, friction error dropped from 0.21 to 0.15 (on a 0 to 1 scale) compared with the earlier network, and stiffness was within about 5 % of its range. On a real wet whiteboard the error for a hind foot fell from 0.33 to 0.02. A replay in simulation confirmed that the predicted values reproduce the real slipping. From only 22 images, the camera network found the floor with 82 % overlap and got the board's friction within 0.03.

  6. Why it's a step forward

    Because the numbers are in the simulator's language, future walking and navigation controllers trained in simulation could take friction and stiffness as inputs and get them from a camera on the real robot. Still open: estimates wobble when the robot stands still, a thin 5 cm foam board went unnoticed, the replay check only works on flat rigid ground, and that anticipating controller is future work.

Words used below
Friction
how grippy the ground is; simulator value 0 (ice) to 1 (grippy)
Stiffness
how hard the ground is; simulator value 1 (soft) to 10 (rigid)
Proprioception
the robot's sense of its own body: joint angles, speeds, tilt
Self-supervised
the robot makes its own training labels, here from its footsteps
Confidence mask
the pixels whose prediction can be trusted
1 / 8
friction stiffness camera / vision
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper (CC BY 4.0); the text is a plain-language walkthrough.

Scene 1

Ground that looks fine

In short: the hazards that matter most for a walking robot are often invisible in the shape of the ground, and today the robot only learns about them by stepping on them.

Legged robots can now walk over very rough terrain, and they are being used for search and rescue, planetary exploration and inspection of hazardous sites. The paper points to a different kind of danger: non-geometric hazards. A water-covered floor, a patch of ice, soft soil or deep snow can look completely flat to a depth camera, yet a foot slips or sinks as soon as it touches them.

Walking controllers trained in simulation are already robust to this: during training the simulator randomly changes how grippy (friction) and how hard (stiffness) the ground is. But the controller only notices these properties through contact, by feeling a foot slide or sink. The authors want the robot to anticipate them from its camera, before the foot lands, the way animals judge ice or mud from experience.

Animated. Two numbers describe how the ground behaves: low friction makes the foot slide, low stiffness makes it sink. Motion illustrative.
Scene 2

The missing link

In short: earlier “learn-from-footsteps” vision methods predicted scores a simulator can't use; this paper predicts the simulator's own friction and stiffness values instead.

Training entirely in simulation doesn't solve this. You can render photorealistic grass, but it is hard to make simulated grass behave exactly like real grass underfoot. So a camera-to-physics link learned in simulation would not transfer.

The opposite route is self-supervised learning on the real robot: the robot walks, measures something about each footstep and uses it to label the camera images. Previous work did this for traversability (how well the robot kept its commanded speed), learned “ground reaction scores”, or the sound the ground makes. These are useful for choosing where to go, but they are signals a simulator can't reproduce accurately, so a controller trained in simulation can't be given them as inputs.

The paper's move is to make the label the simulation parameter itself. Note what this means: a predicted stiffness of 0.5 is “the simulator setting that makes the robot move as it does here”, not a physical stiffness in N/m. That is exactly what a simulation-trained policy needs.

Paper Fig. 1. The two stages. Top: the interaction pipeline learns “how it feels” from body motion and terrain shape. Bottom: its per-foot labels teach the vision pipeline “how it looks”, which then predicts every pixel and marks unconfident areas.
Scene 3

Practice in simulation

In short: in a simulated world of random patches the right answer is always known, so a network can learn to read friction and stiffness off the robot's motion, one value per foot.

The data comes from legged gym (a GPU robot simulator), extended with procedurally generated terrain and a soft-contact model so that feet can actually sink. Every terrain patch gets a random shape, a friction between 0 and 1 and a stiffness between 1 and 10. An existing robust walking controller (the lab's 2022 perceptive locomotion policy) drives ANYmal over them with random commands up to 1.3 m/s forward or sideways and 0.3 rad/s turning.

For every time step the label is the value under each foot. While a foot is in the air, its label is the patch directly below it. In total they recorded 18.3 hours of simulated walking for training and 5 hours for testing. Training takes about 30 minutes on one GPU (batch 64, Adam, learning rate 0.001, 100 epochs). Labels are normalised on the fly and weighted so that rare values count as much as common ones.

  • friction 0–1 per patch
  • stiffness 1–10 per patch
  • 18.3 h train · 5 h test
  • 1 step = 20 ms
Scene 4

The physical decoder

In short: two identical networks (one for friction, one for stiffness) remember the last second of motion and use a learned gate to decide how much to rely on the terrain shape.

Each network sees the same input: the last 50 observations (1 s). One observation contains the robot's proprioception (the walking command, joint positions and velocities, body motion and the phase of each leg) and the terrain geometry, sampled as heights on circles around each foot, the same pattern as the lab's 2022 controller.

Inside, a one-layer GRU (a recurrent “memory” with 100 units) keeps what matters from that history, and self-attention layers refine it. Then the network splits into a proprioception-only path and a proprioception + geometry path, each ending in a small MLP (sizes 64, 32, 4). A sigmoid gate (MLP 64, 64, 4) mixes the two. The output is four numbers, one per foot.

Animated. The authors observed that friction leans on the body-only path and stiffness on the path with terrain shape: a sinking foot shows up in the height samples around it, while slipping shows in the body's motion. The percentages are illustrative; the paper reports the tendency, not values.

Two differences from the 2022 “belief” decoder that inspired it: this one predicts stiffness as well as friction, and it gives a value for each foot rather than one shared estimate. It is also trained on its own rather than as a side task of the walking policy.

Paper Fig. 2. The twin networks, one for friction and one for stiffness, each with memory, attention, two prediction paths and a gate.
Scene 5

Foot by foot

In short: the decoder tells the four feet apart, beats the earlier decoder on friction, estimates stiffness well, and works on the real robot without retraining.

In simulation

When ANYmal walks from a patch with friction 0.95 onto one with 0.05, the front foot's estimate drops first, at step 225, and the hind foot's about 60 steps later, so the network really tracks each foot rather than an average. There is a lag of about 20 steps (0.4 s): a foot has already crossed into the new patch but hasn't yet pressed on it.

Animated. One estimate per foot versus one estimate for the whole robot. Schematic timing.

One honest detail: on grippy ground the estimate sits near 0.6, not 0.95. Between 0.6 and 1.0 nothing slips, so the robot moves the same way and the exact value simply can't be observed from motion. For adapting a walking or navigation policy, “grippy” is enough.

Paper Fig. 6. Simulation tests. (a) Friction: each foot's estimate drops after it steps onto the low-friction patch. (b) Stiffness: estimates follow the change from rigid to soft ground.
5 h simulated test (mean absolute error)Baseline (2022 decoder)Ours
Friction (range 0–1)0.210.15
Stiffness (range 1–10)not estimated0.46

What this means: friction error is about 29 % lower than the earlier decoder, and stiffness is off by roughly 5 % of its range; both methods still can't tell very grippy values apart.

Paper Fig. 7. How the errors are spread over the 5-hour test set, for friction (vs. the baseline) and stiffness.

On the real robot

For friction, ANYmal was driven backwards from the lab floor onto a whiteboard covered with water. The new decoder distinguishes the feet: hind feet on the slippery board drop to about 0.1 while the front feet, still on the floor, stay higher. The baseline gives the same value for all feet. For stiffness, the robot walked over a soft foam board and the estimate dropped on the foam and rose again on the floor.

Real run. ANYmal backs onto the wet whiteboard. When the left hind foot (circled) slips, its estimate from this paper (purple, lower plot) falls to about 0.2, while the earlier friction-only model (grey) stays near 0.6. The upper plot is the left front foot, still on the floor. Footage: Robotic Systems Lab, ETH Zurich, from the paper's supplementary video (cropped to hide subtitles).
Paper Fig. 8. Real robot on a wet whiteboard. The hind foot's estimate (top) drops while it is on the slippery board and comes back afterwards; the baseline hardly separates the phases.
Mean friction error, LH foot (Table I)BaselineOurs
Wet whiteboard0.33 ± 0.110.02 ± 0.08
Lab floor0.00 ± 0.000.00 ± 0.00

What this means: on the slippery board the new decoder is almost always inside the “correct” range found by the digital twin (next section); on the floor both are fine because that range is wide (0.3–1.0).

Paper Fig. 9. Real robot walking over a soft foam board: stiffness estimates dip on the foam and recover on the rigid floor.
Scene 6

Digital twin check

In short: replaying real motion in simulation shows which simulator friction values reproduce reality: 0.0–0.2 for the wet board, 0.3–1.0 for the floor, and the decoder's real-world estimates fall in those ranges.

Showing that estimates go up and down in the right places is not the same as showing the numbers are right. Real ground has no “simulator friction” to compare against. So the authors built a digital twin, a simulated copy of the experiment: they cut the real whiteboard recording into 400 ms snippets, started the simulated robot in exactly the recorded pose (orientation from the onboard IMU, joint positions and velocities), rebuilt the flat floor in simulation, and replayed each snippet with 10 friction values from 0.0 to 0.9.

Three measures compare the replay with reality: joint-position error, body-orientation error, and how far the foot slides between touching down and lifting off. They used 18 whiteboard snippets where the foot really slipped (slips are rare) and 220 floor snippets without slipping. Longer replays drift away from reality, which is why snippets are short.

Digital twin. The same recorded real-world motion (inset) replayed by simulated robots, each on a patch with a different friction value between 0 and 1 (patch shade). The replay whose motion matches the recording best tells which friction value is right. Footage: Robotic Systems Lab, ETH Zurich, from the paper's supplementary video (cropped to hide subtitles).
Animated. Each column is one simulated replay; the shortest bars mark the friction values that reproduce reality. Slip distances and bar heights illustrative; the best ranges are the paper's.
Paper Fig. 10. Replay error against the simulated friction value. (a) Whiteboard: lowest around 0.1. (b) Floor: everything from 0.2–0.3 upward matches, because nothing slips.

What this means: the decoder's real-world estimates (≈ 0.1 on the board, above 0.4 on the floor) are values that make the simulator behave like reality, so they can be used as labels. The check currently only works on flat, rigid ground; the authors assume, but have not yet shown, that stiffness transfers the same way.

Scene 7

Teach the camera

In short: while the robot walks, each footstep's friction and stiffness are painted into camera images taken earlier, and a small image network learns from them live.

The vision side follows the lab's earlier Wild Visual Navigation system. Every camera image goes through DINOv2, a large pre-trained image model that describes each pixel with a 384-number feature vector. The images (scaled and centre-cropped to 910 × 910) are the only visual input; no depth is needed, because friction and stiffness don't depend much on geometry and the features already carry a lot of it.

Three tasks run in parallel. Inference predicts friction and stiffness for new images. Self-supervision keeps the mission graph: it stores an image's features, camera pose and an empty label mask every 1 m the robot travels (0.2 m indoors), and when the decoder labels a new foothold, projects it into every stored image within 5 m. Learning samples stored images that have labels and updates the network, at about 2 updates per second.

Animated. A footstep ahead of an earlier camera position lands on a specific pixel of that image; footsteps farther than 5 m are not used. Schematic.

The image network is tiny: an encoder-decoder MLP with layers 128, 32, 128 and 384 + 2. The two extra outputs are friction and stiffness; the 384 outputs try to rebuild the input feature (used in the next section). The loss is 0.9 × rebuild error + 0.1 × prediction error, trained only on labelled pixels.

Paper Fig. 5. The online training framework: inference, self-supervised labelling through the mission graph, and continuous learning.
Paper Fig. 3. The image network: it rebuilds each pixel feature and predicts friction and stiffness at the same time.
Scene 8

What to trust

In short: pixels the network can rebuild well are trusted; a two-bump fit decides the cut-off per image with nothing to tune, and indoors the result matched the real floor with IoU 0.82.

The robot only steps on a small part of the scene, yet the network outputs a friction value for every pixel, walls and trees included. The rebuild output solves this: the network learns to rebuild features of ground it was taught, so taught-like pixels have low rebuild error and everything else high error. As training goes on, the histogram of errors turns from one bump into two.

The earlier system fitted a single bell curve to the labelled pixels and needed a threshold setting. Here a two-component Gaussian Mixture Model is fitted to each image's error histogram, and the cut sits where the two components meet. It needs no tuning across datasets and adapts per image.

Animated. The two curves are fitted live with expectation–maximisation on made-up pixel errors; watch the threshold settle between the bumps.
Paper Fig. 4. Outdoors: the error histogram goes from one bump before training to two by step 19, and the confident area (yellow-orange) shrinks to the path.
Paper Fig. 11. Indoors, left to right: the image with projected foothold labels, the raw prediction for every pixel, and the prediction with the confidence mask. The whiteboard stands out as low friction.
Indoor test, 22 training images (Table II)Friction errorMask accuracy (IoU)
Wet whiteboard0.03 ± 0.050.82 ± 0.05
Floor0.00 ± 0.01

What this means: from a couple dozen images, the camera alone gives friction values within the digital-twin ranges and marks the walkable floor with 82 % overlap to hand labels; IoU (intersection over union) is the overlap of predicted and true areas.

Paper Fig. 12. On an 82-image outdoor dataset, the two-component fit (purple) reaches high mask accuracy much sooner than the earlier single-Gaussian threshold (yellow).

The supplementary video also shows the pipeline outdoors: friction on snowy paths and stiffness while walking through snow, with the network adapting online.

Outdoor run. On a forest trail, the camera network's friction prediction (right, bottom; darker blue = grippier) is projected onto the robot's map as it walks. Shown qualitatively in the video; the paper reports no outdoor friction numbers. Footage: Robotic Systems Lab, ETH Zurich, from the paper's supplementary video.
Wrap-up

Limits and what's next

In short: the paper builds the camera-to-simulator-physics link; the controllers that would use it are the next step.

  • Stability. When the robot stands still, friction estimates drift. A thin 5 cm foam board was not detected as soft on the real robot. The authors suspect too little training data or noisy real terrain geometry, and suggest adding noise during training.
  • Validation. The digital twin only covers flat, rigid surfaces and needs manual terrain reconstruction; stiffness has not been checked numerically in the real world yet.
  • Scope of the vision tests. Numbers come from an indoor dataset (22 training images) and an 82-image outdoor dataset for the mask; the off-road runs in the video are qualitative.
  • The payoff. The goal is walking and navigation policies, trained in simulation with friction and stiffness as inputs, that receive those values from a camera on the real robot. The paper lays the ground for that; the policies themselves are future work.