Science Robotics · 2024 · an animated walkthrough

Learning robust autonomous navigation and locomotion for wheeled-legged robots

The paper in 102 seconds · narrated · sound on
Transcript

Kilometres of city streets, on its own. It drives on the flat and climbs stairs, slips past people on the path, and fits through doorways. How does one robot learn how to move, and where to go?

City deliveries don’t end at the road. They cross kerbs, climb stairs and go indoors. Wheeled robots stop at the stairs; legged robots climb, but at half a walking pace. Wheels on legs can do both. But the robot must choose, moment by moment, to roll or to step — and no animal moves like this. And at 2 m/s, a planner that thinks for a second is already out of date.

So the team trained two neural networks by reinforcement learning — trial and error in simulation. A navigator sets a target speed ten times a second; a locomotion network turns it into joint and wheel commands. Nobody told the legs when to step. They drive on the flat, and trot up stairs. The navigator practised in worlds built like video-game levels, with safe routes and moving boxes. It remembers where it has been, so when the way is blocked, it backs off and finds the stairs. On the real robot: with its path blocked, it followed the wall and took the stairs.

Against a classic planner on the same legs, only it avoided every collision — deciding in a third of a millisecond. In Zurich it drove itself 8.3 km, three times faster than a legged robot, with a 53% lower cost of transport — the effort per metre travelled. A navigator that knows its own legs: a step toward robots that deliver the last mile.

Footage: Robotic Systems Lab, ETH Zurich — Movie 1 and Movie S2 of the paper; page loops from Movie S2 and Movie S3. Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

A four-legged robot with wheels for feet learned in simulation how to move and where to go, then drove itself for kilometres through Zurich and Seville.

  1. Why this matters

    City deliveries need robots that are quick on pavements but can also handle kerbs, stairs and doorways. Wheeled robots stop at the first staircase. Legged robots climb, but ANYmal averages about 2.2 km/h, half a walking pace, and runs for about an hour.

  2. What makes it hard

    With wheels on its legs, a robot can roll and step, but it must decide which to do at every moment, and no animal moves like that, so there is nothing to copy. At several metres per second it also has to react quickly. A plan that takes seconds is out of date before it arrives.

  3. What people did before

    Earlier controllers switched gaits with hand-written rules, fixed gait sequences or trajectory optimisation, which is slow and needs a good first guess. Navigation stacks drew cost maps and planned paths over several seconds. They assumed the robot would follow the path exactly, which led to pauses, zig-zags and near misses.

  4. What this paper does

    Two neural networks, stacked like a driver and a car. A locomotion policy takes a target speed and works out the gait itself. A navigator looks at the terrain, remembers where it has been and sets that target speed ten times a second. Both learned by trial and error in simulated worlds generated the way games generate levels.

  5. What they showed

    Kilometre-scale autonomous missions in Zurich and Seville, around pedestrians and up stairs. In Zurich it covered 8.3 km at 1.68 m/s on average, about three times a legged robot’s speed, using 53 % less mechanical effort per distance. Against a classic planner it was the only one with no collisions.

  6. Why it’s a step forward

    The navigator knows what its own legs can do, so gaits, speed and route choices fit together without hand-tuned glue. That makes wheeled-legged delivery robots practical. The limits: it sees only about 3 m ahead and understands shapes, not meanings such as “this is a lawn”.

Words used below
Reinforcement learning
learning by trial and error in simulation, guided by a score
Policy
the trained neural network that decides what to do next
Hierarchy
a navigator policy giving speed targets to a locomotion policy
Waypoint
a point on the route the robot should head for next
Cost of transport
energy spent per unit of weight and distance; lower is more efficient
Navigation graph
a network of safe spots and connections, borrowed from video games
1 / 8
navigator (where to go) locomotion (how to move) the full system (ours)
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper (CC BY 4.0); the text is a plain-language walkthrough.

Scene 1

A delivery route is not just a road

In short: a delivery robot has to be fast on long flat stretches and still get over kerbs and up stairs. Wheels alone can’t do the second, and legs alone are too slow for the first.

The paper starts from last-mile delivery. More people live in cities, traffic keeps growing and customers want faster deliveries. The last stretch of a delivery often isn’t even on a street: it goes along pavements, up steps and indoors. A robot for that job has to be quick and efficient on flat ground and still able to climb.

Ordinary wheeled robots cannot get over obstacles like stairs. Legged robots can, but they are slow and short-lived for this job. The authors cite their own lab’s ANYmal: at most about 1 hour of operation, and on average about 2.2 km/h, half a person’s walking speed.

A wheeled-legged robot puts driven wheels at the ends of its legs: 12 leg joints plus 4 wheels. It can cruise like a vehicle and step like an animal. The rest of the paper is about getting such a machine to do this by itself over kilometres of real city, not just in a lab.

Paper Fig. 1. Where the system was tested: locomotion challenges (stairs, steps, rough ground) and navigation challenges (people, obstacles, narrow spaces), at locations in Zurich and Seville.
Scene 2

Roll or step, and decide fast

In short: the robot has to invent its own mix of rolling and stepping, and at its speeds it has to decide much faster than classic planners can.

Problem 1: when to roll and when to step

Most legged-robot controllers borrow gait patterns from animals, but no animal has wheels, so there is nothing to borrow. The best gait depends on direction, speed and terrain; rolling as much as possible saves energy. Earlier wheeled-legged controllers used simple rules for when to step, or pre-defined gait sequences. Trajectory optimisation can discover rich motions such as terrain-aware gaits and skidding, but it is computationally expensive, needs a near-optimal starting guess, and faster versions simplify the model (for example ignoring wheel dynamics), which hurts on the real robot.

Problem 2: navigation that doesn’t know the body

Classic navigation estimates a cost map of how passable the ground looks, then samples a path through it. Such planners ignore how the robot actually moves: how well it tracks commands depends on terrain, speed and gait. The results are frequent turning and stepping. They are also slow: state-of-the-art sampling planners for legged robots typically take several seconds per path, far too long at several metres per second around people.

Finally, full autonomy stacks glue many separately built modules together with heuristics. In the DARPA Subterranean Challenge, the authors’ team Cerberus saw robots pause mid-path to re-plan and zig-zag while trying to follow a route.

Animated. At 2 m/s, one second of planning means 2 m driven on an old decision; a policy that decides every 0.1 s updates after 0.2 m.
Scene 3

Two neural networks, stacked

In short: a navigator network decides where to go and sends a target speed; a locomotion network decides how to move the joints and wheels to achieve it. Because the navigator was trained with these exact legs, it only asks for what they can do.

The global route comes from a map, as a chain of graph nodes. From it the system picks two nearby waypoints, WP1 and WP2, a little like a driver looking at the next two turns. The navigation policy (the paper calls it the high-level controller) turns those into a velocity command: forward, sideways and turn rate, 10 times per second, matching the update rate of the onboard terrain map. The locomotion policy (low-level controller) turns the command into 12 joint position targets and 4 wheel speed targets, 50 times per second.

The navigator replaces what used to be separate path-planning, path-following and messaging layers. Its inputs:

  • Height map sampled from the robot-centric elevation map, 3 m to the front and 1.5 m in other directions, plus the scans from 0.1 s and 0.2 s earlier so it can see things move.
  • The locomotion policy’s hidden state: its internal memory, which encodes things like terrain properties and disturbances, used instead of estimated velocities or gravity direction.
  • Position memory: the last 20 visited positions, recorded every 0.5 m (about 10 m in total), with how long it stayed at each.
  • Waypoints: the two current waypoints, the two previous ones and its own last three commands, which help it drive smoothly.

Its output is drawn from a Beta distribution, which can only produce values inside hard limits: forward −1.0 to 2.0 m/s, sideways ±0.75 m/s, turning ±1.25 rad/s. The forward range is shifted so the robot tends to face where it drives, which is where its camera looks.

Animated. As the average command moves, the Beta distribution (filled) never leaves the allowed range; a Gaussian with the same spread (dashed) would sometimes ask for speeds outside it (red).

During training, the navigator’s reward includes the locomotion policy’s own reward terms (tracking, smoothness, low torque), scaled to a similar size. That is what makes it mobility-aware: it earns more by asking for smooth, low-effort motions the legs can actually follow. Separating the two levels also let the team develop them independently and reuse a trained locomotion policy for other tasks.

Design (Table S1)SPL, 5–10 mSPL, 10–20 m
Ours: navigator at 10 Hz over locomotion at 50 Hz0.8970.689
Navigator also at 50 Hz (no slower top level)0.7980.370
One end-to-end policy straight to the joints0.3040.045

What this means: splitting the job in two, with the navigator thinking at a slower pace, is what makes long routes solvable; a single network doing everything almost never reached distant goals. (SPL is success weighted by how close the path was to the shortest; 1.0 is perfect. 1,000 random simulated worlds.)

Paper Fig. 2. The robot and its payload (three LiDARs, a stereo camera, a delivery box, a 5G router and GPS), the navigator-over-locomotion structure, and the training world with pre-generated obstacle-free paths.
Scene 4

The legs pick their own gait

In short: nobody programmed when to step. The locomotion policy was trained to follow speed commands with little effort, and it discovered by itself to drive on the flat, trot on stairs and creep over big obstacles.

The locomotion policy builds on the lab’s earlier perceptive walking controller (Miki et al., 2022) with two notable removals. First, there are no built-in gait rhythms: the earlier controller used a central pattern generator, a built-in stepping rhythm, which is removed here, so the network outputs joint positions and wheel speeds directly. Second, there is no state estimator. Instead of estimated body velocity and orientation, it reads raw accelerometer and gyro signals from the IMU and the joint encoders. Estimators often go badly wrong when wheels slip or the robot drops down a step; the paper’s Movie S4 shows such a failure.

It was trained in two stages, a technique called privileged learning. A teacher policy learns with reinforcement learning in simulation while seeing things a real robot can’t measure: noise-free joint states, foot contacts and contact forces, terrain normals, true body velocity and gravity direction. Then a student policy, the one that runs on the robot, learns to imitate the teacher using only noisy, biased sensor data. It has a recurrent memory (a GRU) that builds up a belief about the world over time.

Animated. The teacher trains with simulator-only information; the student then learns to copy its actions from sensors a real robot has.

Training commands were sampled from forward ±2.5 m/s, sideways ±1.2 m/s and turning ±1.5 rad/s. The rewards favour tracking the command and penalise torque, jerky motion and body contacts with the ground. Rolling uses the least effort on easy ground, so the policy rolls there and steps only where it must. The behaviours in Fig. 6:

  • an asymmetric creep-and-drive gait over a large discrete obstacle;
  • a trot like a point-foot quadruped on stairs and steep hills;
  • driving over bumps about the size of the wheel radius, with the legs acting as active suspension;
  • a lower body when going down a slope, to avoid tipping;
  • a controlled drop off a table about 60 cm high, with a full flight phase;
  • crossing a block about 40 cm high with all wheels in the air for a moment, crawling forward on its knees.
Real run. The locomotion policy taking the robot over a concrete block (the movie labels steps up to 40 cm): it lifts its legs one at a time, climbs on and drives off the far side. No gait sequence was programmed for this. Footage: Robotic Systems Lab, ETH Zurich (paper Movie S3).

On flat ground it reached 5.0 m/s, 18 km/h. The hardware limit is 6.3 m/s: a joint speed of 45 rad/s times the 0.14 m wheel radius. In a simulated slope test (friction 0.7), stepping appeared only on steep slopes with commands above 0.5 m/s, and stepping let it climb steeper slopes. It also handled higher steps going down than going up. The lab’s previous model-based controller could not operate in these environments.

Paper Fig. 6. Gaits that emerged on different real terrains (with wheel-contact timelines), the table drop and the high block, and the largest step heights and slopes crossed at each commanded speed, going up versus down.
Paper Fig. S1. Driving toward a waypoint behind a 1 m × 1 m obstacle. The hierarchical controllers step only when they meet the obstacle; a baseline with a fixed stepping rhythm (1 Hz) keeps stepping even when it doesn’t need to.
Scene 5

Game-style practice worlds

In short: the navigator practised in endless generated worlds that always have a solvable route, borrowed from how video games let characters find their way.

With randomly scattered obstacles and random goals, many goals are unreachable, and a learner mostly learns to give up. The authors borrowed the navigation graph from game engines such as CryEngine and Unreal: a network of safe spots and connections that comes with the level.

To generate levels they used Wave Function Collapse, a procedural-generation algorithm that assembles new layouts from example tiles so that neighbouring tiles always fit. There are three tile types, Stair, Floor 0 and Floor 1, with rules such as “stairs along x connect only to floors along x.” The output is both the terrain and the connectivity graph. The stair and floor parameters come from the ranges the locomotion policy could handle in its own training, so every route is feasible for the legs.

Animated. Cells are fixed one at a time, each respecting its neighbours; when the grid is complete, the graph of safe connections comes for free.

Each episode, Dijkstra’s shortest-path algorithm connects two random nodes. Two waypoints are placed along the route with a look-ahead distance drawn from 5–20 m; the final node is given twice at the end. Early in training the navigator is rewarded for speed toward the nearest waypoint along the route, which follows detours and tight turns. This dense reward is gradually reduced until only reaching waypoints counts. Boxes appear at random and drive toward the robot at 0.1–0.5 m/s, to practise dodging.

Navigator trained…5–10 m: SPL (success)10–20 m: SPL (success)
as in the paper (graph paths + WFC worlds + memory)0.897 (0.901)0.689 (0.763)
with random goals instead of graph paths0.858 (0.840)0.497 (0.559)
on rough terrain with random obstacles instead of WFC worlds0.865 (0.871)0.302 (0.305)
without position memory0.873 (0.897)0.526 (0.573)

What this means: on short routes the variants are similar, but for distant goals the structured worlds matter most: without WFC worlds, success on 10–20 m routes fell from 76 % to 31 %. (Table S1, 1,000 random worlds; success = within 50 cm of the goal within 60 s.)

Paper Fig. 8. From the locomotion policy’s training terrains, to tile maps and connectivity graphs from Wave Function Collapse, to a height-map world with a sampled path and moving boxes (white).
Scene 6

A memory of where it has been

In short: the navigator remembers its last 20 positions and was trained to dislike hanging around them, so when a route is blocked it explores instead of repeating itself.

The navigator only sees a few metres around it, so a blocked route is a puzzle. Instead of a general-purpose recurrent memory, the authors gave it a simple, readable one: a buffer of the last 20 positions, stored every 0.5 m, each with how many time steps the robot spent there. A small network in the style of PointNet (1-D convolutions and max-pooling) reads the set of positions regardless of order.

During training an exploration bonus, really a penalty, discourages loitering. Unless the robot is already within 0.75 m of its waypoint, every stored position within 1 m of the robot subtracts the number of steps it spent there.

C = 0  if |probot − wp1| < 0.75 m;   −nibuf  if |probot − pibuf| < 1.0 m  // summed over the 20 stored positions
Animated. When the robot loops back, the remembered spots inside the 1 m circle light up and the penalty grows. This is the pressure that teaches it to try new places.

On the real robot (Fig. 5, Movie S2) this produced behaviours that look deliberate:

  • Blocked path: it reversed, followed the wall looking for an opening, and found the stairs to the waypoint.
  • Narrow passage: it went through two doors with a person standing between them, where the gap was as wide as the robot, without collision (human detection was off in that run).
  • Complex obstacle: with small stairs on one side and a step of 0–50 cm on the other, it either found the stairs or drove along the step until it found a part about 20 cm high to climb.
  • Up versus down: it would cross higher steps going down than going up, an asymmetry that cost maps, which rate ground the same in every direction, usually ignore.
  • People: a camera detector marks each person with a raised patch of 50 cm radius in the height map; the navigator kept its distance from a walking person and then overtook safely.
Real run. The end of the blocked-path test: after backing away from the blocked route, the robot follows the wall, finds the stairs and climbs them toward its waypoint. Footage: Robotic Systems Lab, ETH Zurich (paper Movie S2).
Paper Fig. 5. Local navigation on the real robot: exploring around a blockage, a door-width gap, two solutions to a stairs-and-step obstacle, direction-dependent step heights, and the height-map buffer around a detected person.
Scene 7

Head-to-head with a classic planner

In short: with the same legs underneath, the learned navigator failed less, never collided, decided thousands of times faster and followed its own commands more accurately than a proven sampling-based planner.

The baseline is the sampling-based planner of Wellhausen and Hutter, which team Cerberus used in the DARPA Subterranean Challenge; it was designed for ordinary legged robots. Both methods drove the same locomotion policy. The site from Fig. 5A (start and goal 5 m on either side of a wall-and-stairs obstacle) was laser-scanned and rebuilt in simulation. Each method ran 10 times with start headings between −45° and 45°. Both maps were limited to about 3.5 m, because a larger map slowed the elevation mapping at speeds up to 2 m/s.

Fig. 7E (10 trials each)OursOurs, no memoryClassic planner
Failed to reach the goal30 %80 %50 %
Collided0 %0 %100 %
Compute per decision0.34 ms (mean)—sometimes > 1 s
Mean speed-tracking error0.24 m/s—0.45 m/s

What this means: the learned navigator was the only one with no collisions. It explored beyond what it could see, and it asked only for motions the legs could follow. Without memory, the same policy got stuck in loops most of the time.

The failure modes are instructive. The planner was overconfident about hidden areas and slow to re-plan when they turned out to be blocked. It also assumed the legs would track its path perfectly: far-away pose targets turned into high speed commands and overshoot, so it hit obstacles. Its tracking-error histogram has a second peak at high error, from sudden command changes or commands too close to obstacles that the locomotion policy refused. Ours sometimes stopped exploring after roaming a wide open area, because its 20-position memory was full. Timing was measured on a desktop PC (AMD Ryzen 9 3950X, RTX 2080).

Paper Fig. 7. The point-goal test: setup, each method’s field of view, successful trajectories (ours found the stairs by exploring in either direction), failure cases, and failure rate, collision rate, planning time and tracking-error histograms.
Scene 8

Kilometres in the city

In short: on mock-up delivery missions in Zurich the robot drove itself 8.3 km, three times faster than a legged robot and with about half the mechanical effort per distance.

Getting ready: a digital twin

Before a mission, someone walked the area with a handheld laser scanner: about 90 minutes for a 245 m × 345 m district of Glattpark, Zurich. The point cloud was georeferenced and turned into a mesh. On it a person drew a sparse navigation graph, which also encodes social preferences such as staying off landscaping and private property. On the robot, LiDAR scans are matched to that point cloud to know where it is; the authors found this more robust among high-rise buildings than GPS. A goal is sent over the mobile network, the robot computes the shortest route on the graph (Dijkstra), and an “anchor pursuit” rule picks waypoints about 3 m ahead without cutting corners at graph nodes. The point cloud is used only for localisation. Obstacles come from the live onboard height map.

The missions

The robot visited 13 distant goal points in several runs of over 30 minutes each, 8.3 km in total, with minimal human intervention. It avoided pedestrians (also on slopes and stairs) and thin poles, and it went up steps and stairs and over grass, sand and gravel. It drove on flat ground and switched to stepping on uneven ground.

While moving (Fig. 3B-iii)Wheeled-legged (ours)ANYmal-C (legged)
Average speed1.68 m/s0.55 m/s
Mechanical cost of transport0.160.34

What this means: about three times as fast for about half the mechanical effort per metre (53 % lower). The ANYmal numbers are from a robot that mostly crossed flat urban ground during the DARPA Subterranean Challenge. The text gives 0.16; Fig. 3 marks 0.17.

Mechanical cost of transport here is the positive mechanical power of all joints divided by weight × horizontal speed. It counts only output power: heat and transmission losses are not included. Most of the gain comes from driving: weight spreads evenly over four legs and the leg joints stay nearly still, contributing a cost of transport of about 0.01. The wheels put out about 1.2× the total mechanical power of a walking ANYmal while moving 3.4× faster. The sum of squared leg-joint torques, which tracks heat loss, was 16 % lower even though the robot is about 12 kg heavier.

Paper Fig. 3. The workflow (scan, process, draw the graph, plan and follow), where the robot went and how often, the human-designed graph, and speed and cost-of-transport histograms against ANYmal-C. The aerial image is credited in the figure to CNES, Spot Image, swisstopo and NPOC.

When people stepped in

Three kinds of situation (Fig. 4B). Children in the path: the team stopped the robot proactively, although it had passed adults safely. Tall grass had grown over a trail since the graph was made: the robot stopped in front of it, and the team triggered a re-plan. Lost localisation in geometrically featureless places such as long corridors: the robot stayed safe using its local map but could not reach the goal until localisation came back.

Paper Fig. 4. Things to avoid (people, poles) and things to cross (stairs, steps), and the three situations where the team intervened.
Limits

What is still missing

In short: the robot sees shapes but not meanings, and not very far. Preparing a new district still takes human work.

  • Little semantic understanding. Navigation is geometric; the only semantic input is person detection. Recognising pavements or judging terrain from images would allow better choices.
  • Short, slow perception. The navigator sees at most about 3 m ahead because it relies on elevation mapping. The hardware can move at about 6.2–6.3 m/s (the paper gives both figures), but mapping delays kept autonomous runs well below that. Using raw sensor streams instead of a map is a suggested next step.
  • Human-made maps. Each area needs a scan and a hand-drawn navigation graph.
  • Limited memory. Twenty positions (about 10 m) is enough for local puzzles, but after a wide exploration the buffer fills up and the robot can get stuck (Fig. 7D-i).
  • Training separately. Alternating further training of both levels together gave only marginal gains, so it was not used.