arXiv · 2022 · an animated walkthrough

Team CERBERUS Wins the DARPA Subterranean Challenge: Technical Overview and Lessons Learned

All 36 authors

Marco Tranzatto, Mihir Dharmadhikari, Lukas Bernreiter, Marco Camurri, Shehryar Khattak, Frank Mascarich, Patrick Pfreundschuh, David Wisth, Samuel Zimmermann, Mihir Kulkarni, Victor Reijgwart, Benoit Casseau, Timon Homberger, Paolo De Petris, Lionel Ott, Wayne Tubby, Gabriel Waibel, Huan Nguyen, Cesar Cadena, Russell Buchanan, Lorenz Wellhausen, Nikhil Khedekar, Olov Andersson, Lintong Zhang, Takahiro Miki, Tung Dang, Matias Mattamala, Markus Montenegro, Konrad Meyer, Xiangyu Wu, Adrien Briod, Mark Mueller, Maurice Fallon, Roland Siegwart, Marco Hutter, Kostas Alexis

The paper in 108 seconds · narrated · sound on
Transcript

One hour in a dark underground course. Four walking robots, searching on their own, each building its own map. The score climbs to twenty-three — enough to win the DARPA Subterranean Challenge.

When a mine collapses, rescuers go in blind. Robots could go first, but underground there is no GPS, no light, and radio dies behind the first corner. DARPA made it a contest: one hour, forty hidden objects, and one person to command every robot.

Team CERBERUS sent four-legged ANYmal robots, backed by drones and a rover that carries the network. Each robot tracks itself with layered sensors, so when smoke blinds its cameras, its laser and legs take over. It keeps choosing the route that reveals the most unknown space, and backtracks from dead ends. A walking controller learned in simulation reads the ground, and a planner steers around risky rails.

Carriers drop radio relays. Explorers leave radio range on a time budget, then return to upload what they found. A neural network spots objects in camera images; rays cast into the map place them. One person approves every report — and a gas sensor set too sensitive flooded him with thirty in five minutes.

In the final, CERBERUS tied CSIRO Data61 at twenty-three points, and won by reporting its last point earlier. Four ANYmals walked up to 687 metres each through the course, with no falls. Their lessons: test in the field every month, and never overload the one human.

Footage: Team CERBERUS, from the team’s two-minute Prize Round video (course-camera views from the DARPA Final Event broadcast) and the Prize Round video linked in the paper (robot camera and map views, in the hook and the “687 m” shot), both on the CERBERUS SubT channel. Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

In September 2021, four walking robots, helped by a rover and guided by one person standing outside the entrance, explored a dark, smoky underground course for an hour, found 23 hidden objects and won the DARPA Subterranean Challenge. This paper explains how the whole system worked, and what the team learned the hard way.

  1. Why this matters

    When a mine collapses, an earthquake wrecks a building’s basement, or someone is lost in a cave, rescuers have to go in blind. Robots could go first and report what is inside: where the people are, where the gas is, which way is passable.

  2. What makes it hard

    Underground there is no GPS, no light, often smoke and dust, and the floor turns into stairs, rubble and rails. Radio dies behind the first bend, so a robot is often on its own. The contest also allowed only one person to command every robot, like one air-traffic controller for a whole fleet.

  3. What people did before

    Earlier underground robots each solved one piece: a mine-inspection robot, a drone that flies down the centre of a tunnel, ways to squeeze a 3-D map through a thin radio link. DARPA’s challenge asked for everything at once (moving, sensing, deciding and communicating) in tunnels, basements and caves alike.

  4. What this paper does

    It describes CERBERUS, a “system-of-systems”: four-legged ANYmal robots, small collision-proof drones and a rover on a fibre cable. Each robot can map, plan and search on its own, and they share a radio network they extend by dropping relays. One supervisor decides what each robot should do; the robot works out how.

  5. What they showed

    In the final run CERBERUS scored 23 points, tied with CSIRO Data61, and won on the tie-break by reporting its last point more than a minute before the end. Four ANYmals walked between 240 m and 687 m each, with no falls, and kept their position errors to a few centimetres per step.

  6. Why it’s a step forward

    In 2019 only 2 of 11 teams relied on legged robots; by the 2021 final all top six did. The paper is also honest about what went wrong: an overloaded supervisor, a gas sensor set too sensitive, relays left unused, and robots trapped by obstacles that moved. It ends with lessons for anyone building field robots.

Words used below
Artifact
one of the 40 objects to find: a mannequin “survivor”, backpack, phone, gas, and so on.
SLAM
working out where you are while drawing the map at the same time.
LiDAR
a spinning laser that measures distances to build a 3-D picture.
Elevation map
a grid of ground heights around a walking robot.
Mesh network
radios that relay messages hop by hop; “breadcrumbs” are relays dropped on the way.
Supervised autonomy
robots decide how to do a task; one human decides which task.
1 / 8
robots on their own radio network human supervisor objects to find
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough.

Scene 1

The challenge

In short: DARPA built underground places that are as hard as real disasters, and asked teams to send robots in first, with only one person in charge.

The Subterranean (SubT) Challenge was a three-year, $82 million competition run by DARPA, the US defence research agency. Its goal was plain: give rescuers and soldiers a quick picture of an unknown, dangerous underground space before anyone enters, whether that is a collapsed mine, a basement after an earthquake or a cave system. More than 300 people in 20 teams from 11 countries took part.

Robots were tested in four events: a mine Tunnel circuit (August 2019), an Urban circuit in multi-storey underground buildings (February 2020), a Cave circuit (August 2020, cancelled), and the Final Event (September 2021), which combined all three into one course. DARPA named four areas where robots had to improve:

  • Mobility: narrow passages, sharp turns, drops, stairs, mud, sand, water and debris that falls when you pass.
  • Perception: sensing through darkness, dust, fog and smoke.
  • Autonomy: exploring and mapping with little or no human help.
  • Networking: getting information out when radio signals can’t pass through rock.

Points came from artifacts: mannequin survivors, phones, backpacks, drills, fire extinguishers, vents, helmets, ropes, gas-filled rooms and a “SubT cube”. In the final, each team had 60 minutes to find as many of 40 as it could. A report scored only if it named the right type of object and placed it within 5 m of the true spot, and a team had only 45 tries. Just one team member, the Human Supervisor, could talk to the robots once they were inside; four “pit crew” members could only set up and prepare the robots before the start.

Paper Fig. 1. The CERBERUS robots: ANYmal C walkers in cave and urban sections (top), and the collision-tolerant RMF-Owl, a DJI-based Aerial Scout, the Kolibri flyer and the tethered rover (bottom).

Where earlier work stopped

Underground robotics is not new. The paper points to a coal-mine inspection robot, a tunnel-inspection platform with an arm, a robot that explored mines using a simplified “topological” map of junctions and corridors, drones that steered by the dark centre of a tunnel or flew along dam pipes, a method for sending compact maps over thin links, and a ground–air pair where the ground robot mapped coarsely and the drone filled in detail. Each solved part of the problem. The SubT rules ruled out the usual shortcuts: fragile autonomy, constant teleoperation, specialising in one kind of underground space, falling apart when sensors are degraded, or depending on a radio link that is always there.

The other finalists took different routes: CSIRO Data61 used tracked robots, Spot legged robots and Hovermap drones with their Wildcat mapping; CoSTAR built the NEBULA framework, which plans with an explicit estimate of its own uncertainty; Explorer combined wheeled robots, drones and a Spot; MARBLE ran Spot and Husky robots; CTU-CRAS-NORLAB fielded one of the most varied fleets; Coordinated Robotics and Robotika brought wheeled and flying robots.

Scene 2

The team of robots

In short: legs for rough ground and long missions, wings for places legs can’t reach, and a rover to carry the network; every robot able to work alone.

The name stands for “CollaborativE walking & flying RoBots for autonomous ExploRation in Underground Settings”, after the three-headed dog guarding the underworld in Greek myth. The team joined the University of Nevada, Reno, ETH Zurich, NTNU, UC Berkeley, the University of Oxford, Flyability and Sierra Nevada Corporation.

ANYmal C SubT: the workhorse

For the final the team switched from the older ANYmal B300 to the ANYmal C100 from ANYbotics, a water- and dust-proof (IP67) quadruped, and customised it. Each weighed 55 kg with payload and could walk continuously for 80 minutes. All carried a Velodyne LiDAR, an Alphasense camera unit (three black-and-white and four colour cameras looking front, left, right and up, lit by four high-power LEDs), a CO₂ sensor, a Bluetooth sensor and an NVIDIA Jetson AGX Xavier graphics computer. They came in two roles:

  • Explorers went deep to find objects. They had two dome-shaped LiDARs for terrain mapping, which cope with dust, water and shiny surfaces, and a pan-tilt inspection head with a 10× zoom camera, a thermal camera, a microphone and a spotlight.
  • Carriers used four lighter depth cameras instead, which freed room for four radio “breadcrumbs” held by electromagnets. To drop one, the robot tilts its body and switches off the magnet; the relay slides out on carbon tubes and lands upright.
Paper Fig. 2a. Explorer, with the inspection head on top.
Paper Fig. 2b. Carrier, with the relay magazine on its back.

Flying robots and a rover

For narrow shafts and stairwells the team built RMF-Owl, a 1.46 kg collision-tolerant drone 38 × 38 × 24 cm in size, with a LiDAR, a camera and LEDs, running the whole autonomy stack on a small board computer; it flies for about 10 minutes. The larger Kolibri (6 kg, 22 min, a multi-directional tricopter) was built but not deployed because the final course was so narrow. A roving robot (135 kg) reeled out a 300 m fibre-optic cable from the base station and carried a radio with a directional antenna; it was driven by the supervisor and parked at a junction as a network hub.

Paper Fig. 5b. The roving robot: a cable reel at the back, a radio panel antenna at the front, and cameras and LiDAR so it can also work as a backup explorer.
Scene 3

Where am I?

In short: every robot tracks its own position with several sensors stacked in layers, so when one goes blind the next one takes over.

The same camera unit, lenses, timing and lights went on every legged robot and on Kolibri. That made tuning simpler and helped the software recognise places seen by another robot. Getting sensors to agree took care: the team calibrated cameras with a large grid of AprilTag markers (6 × 6 tags about 10 cm wide) and found that automatic camera-to-LiDAR calibration didn’t work with a sparse 16-beam LiDAR, so they built a tool to align the two by hand.

CompSLAM: layers of fallbacks

Each robot ran CompSLAM on board, with no need for the base station. It works like a relay race in which each runner checks the baton before taking it:

  1. Cameras + thermal camera + motion sensor (IMU) estimate motion first. Thermal images still show structure in smoke, and sparse LiDAR points help set the scale of what the cameras see.
  2. LiDAR odometry matches each new laser scan to the previous one. Before using the camera estimate as a starting guess it checks that it is plausible (within the robot’s possible speed, with uncertainty not growing too fast). If not, it falls back, in order, on leg odometry (ANYmal counting its own steps), the IMU alone, or “keep the same speed”.
  3. If the scan match itself is poorly constrained (a long smooth corridor looks the same everywhere along its length), the uncertain directions are taken from the starting guess instead.
  4. LiDAR mapping matches scans to the growing map, with the same health check.
Animated. When smoke fails the camera stage’s health check, the LiDAR stage takes its starting guess from leg odometry instead. Health values are illustrative.
Paper Fig. 6. CompSLAM’s layers: each estimate flows downward and is refined, with fallbacks when a stage reports poor health.
Real run. ANYmal 4 in the urban section of the Prize Round, sped up 20×: its four cameras (top) and the 3-D map and path that grow as it walks (bottom). Footage: Team CERBERUS, Prize Round video linked in the paper.

A shared map at the base: M3RM

When in radio range, robots also sent compressed map chunks (about 2 MB each) to M3RM, a server at the base that merged all robots’ maps and looked for places seen twice (“loop closures”) to cancel drift. It worked well afterwards but struggled during the run: a long, white, texture-free corridor at the urban entrance gave its camera features almost nothing to hold on to, so its maps skewed and it slowed down. The supervisor therefore mostly relied on each robot’s own CompSLAM estimate.

RobotCompSLAM, during runCompSLAM, re-runM3RM, re-run
Step-to-step error (relative pose error, translation, m)
ANYmal 10.070.03–
ANYmal 20.050.04–
ANYmal 30.050.04–
ANYmal 40.060.04–
Overall drift (absolute pose error, translation, m)
ANYmal 10.600.720.25
ANYmal 21.811.290.36
ANYmal 30.220.230.20
ANYmal 41.241.000.24

What this means: on the day, each robot’s own estimate drifted by 0.2–1.8 m over the whole run, well inside the 5 m scoring radius. Re-run afterwards with more time, the merged multi-robot map cut that to 0.2–0.4 m. (The “re-run” CompSLAM restores a finer point-cloud setting the team had coarsened as a safety margin; step errors from M3RM can’t be computed because it keeps only key frames.)

Scene 4

Where next?

In short: one exploration planner on every robot keeps asking “which nearby route shows me the most new space?”, and remembers unexplored openings for later.

All CERBERUS robots, walking and flying, used the same graph-based exploration planner, GBPlanner2, with only robot-specific settings changed. It works on a 3-D grid map of free and occupied space (built with Voxblox) and has two layers:

  • Local: inside a box around the robot, whose size adapts to the space, it scatters points and links them into a graph of collision-free moves. For walking robots each move must also have ground under it and a slope the robot can manage. It finds the shortest route to every point and picks the one that would reveal the most unknown volume.
  • Global: a sparse graph spanning everything explored so far keeps “frontier” points that still promised new space. When the local layer finds nothing worth doing (a dead end), the global layer sends the robot back to a frontier. At every step it also checks whether the battery still allows one more round and the trip home; if not, it heads home.

Two details came straight from field testing. “Hanging vertices” keep points over gaps in the map (negative obstacles such as a drop the sensors saw only partly) without ever sending the robot there, and before the final the planner was tuned to squeeze through openings as small as 1 m × 1 m, which the final course had plenty of. After the second preliminary run showed routes over slopes that were too steep, the global layer was changed to compute several shortest routes and take the one whose steepest part is gentlest.

Animated. Three routes out of a junction, scored by how much unknown space each would reveal. Numbers are illustrative.
Paper Fig. 10. GBPlanner2’s two layers: a dense local graph for exploring nearby and a sparse global graph for repositioning and going home, plus the ground check used for legged robots.
Paper Fig. 28d. Global repositioning in the Prize Round: after finishing part of the cave, the supervisor picked a goal at the start of the tunnel and the planner routed the robot there along its global graph.
Scene 5

Every step

In short: a height map of the ground feeds both a learned walking controller and a planner that weighs the risk of a misstep five times more than lost time.

The elevation map

An elevation map stores one ground height per grid cell. ANYmal builds it around itself on its GPU from its depth sensors and odometry: 8 m × 8 m at 4 cm cells, updated at 20 Hz on Explorers and 45 Hz on Carriers. It also clears “ghost” obstacles by tracing rays from the sensor to each measurement: if a ray passes below a stored height, that height was wrong.

Paper Fig. 7. ANYmal’s elevation map. Coloured dots are the heights the walking controller reads around each foot; the black-and-white layer marks where a foot can safely go.

A walking controller learned in simulation

The legs were driven by a neural-network controller trained with reinforcement learning in simulation. It reads the terrain heights in circles around each foot, and during training it met random terrains, varied robot mass and ground friction, pushes, and deliberately corrupted maps, so it learned to rely on its sense of touch when the map is wrong. One controller handled the whole course (tunnel, urban stairs and cave) with no need to tell it which gait to use. It also learned to tilt the body to drop relays on slopes. (That controller is described in its own paper, Miki et al., Science Robotics 2022.)

Real run. An ANYmal crossing the rocky cave section of the Prize Round, the kind of ground the learned controller handled with no falls. A few seconds of course-camera shots; score overlays cropped. Footage: Team CERBERUS, two-minute Prize Round video (course cameras: DARPA SubT Final Event).

A navigation planner that respects risk

The ANYmal navigation planner refines the exploration path locally. It models the robot as a torso that must stay clear of obstacles and four “reach” boxes (one per leg) that must touch steppable ground. A tiny network of only 120 parameters, trained on just 20 hand-labelled height maps, marks where feet can go. A second learned model predicts the risk, energy and time of a move from the height map, using data from a walking policy on random simulated terrain. Energy mattered little (the batteries outlasted the run), but one navigation failure could end the mission, so the cost was risk plus time, with risk weighted five times higher.

Paper Fig. 29c. In the Prize Round the learned cost led ANYmal around high rails; when a rail had to be crossed, the robot turned first so it could walk straight along it.

A behaviour tree (a flowchart of checks that runs over and over) coordinated all this: a watchdog made the robot retrace its steps if it stopped moving, the supervisor’s commands took priority, a timer called the robot home when its budget ran out, and the height map was reset automatically if the robot’s pose ever looked impossible, for example when drift made the map rise to the robot’s knees.

Scene 6

Staying in touch

In short: the team carried its own radio network into the ground, and designed robots to keep working, and come back, when that network ran out.

The backbone was a 5.8 GHz mesh network of commercial Rajant radios: the base station, every ANYmal and the rover were mesh nodes that connect to as many neighbours as they can and route data between them; each also broadcast a normal Wi-Fi network for the drones. The base had two links: a radio in the staging area, and the rover’s 300 m fibre cable. Carriers could add up to eight breadcrumbs (four each), each weighing 245 g with a 2-hour battery and an antenna that unfolds 17 cm upright after landing. Where to drop them was left to the supervisor, since radio underground is so unpredictable.

Paper Fig. 17b. A Carrier lowers its body and releases a breadcrumb relay.

On top of the radios, the software (ROS 1 with the Nimbro Network package) split data into essential streams (heartbeat, telemetry, object reports, map data) that every robot always sent, and heavy non-essential ones such as camera video, which only the robot the supervisor had selected could send. That kept the shared bandwidth free.

Working out of range

The supervisor could let a robot leave coverage with a time budget. It explored for that long, then walked back to where it started, so a task took about twice the budget. Without a budget, a “Wi-Fi safety” check made the robot backtrack as soon as its pings to the base kept failing. In the Prize Round budgets ran from 5 to 13 minutes. The team’s view: short, frequent trips, with the supervisor checking in often, paid off compared with sending robots on long missions and waiting for them until the end.

Animated. A budget of B minutes means about 2B minutes out of contact: B exploring, B walking back.
Paper Fig. 16. The network in the final: mesh nodes (base, rover, ANYmals, breadcrumbs) form the backbone; drones join as ordinary Wi-Fi clients.
Scene 7

Finding things

In short: robots spot objects themselves, turn a glimpse into a 3-D position, and only report when they are confident; the human makes the final call.

Each robot ran an object detector (a YOLOv3-style neural network) on its camera streams: 3 times a second per camera on ANYmal, next to the elevation mapping on the same Jetson. It was trained on 40,007 hand-labelled images collected in basements, tunnels, mines and caves, many lit only by the robots’ own LEDs, plus 384 images of scenes that had fooled it before. The team hunted deliberately for failure cases: look-alike background colours, motion blur, smoke, far-away objects.

ObjectLabelled images
Survivor6,477
Helmet6,053
Fire extinguisher5,357
Vent5,350
Rope5,219
Drill4,580
Cell phone3,556
Backpack3,415
Total40,007

What this means: the classes were roughly balanced. Phones were labelled too, but in practice they were found by their Bluetooth signal.

From a box in an image to a point on the map

For each pixel inside a detection box, the robot casts a ray into its 3-D map; the median of the hit points is the object’s position. A small sphere around it collects later sightings of the same object, the centre is averaged, and a Bayes filter (a running tally of evidence) raises its confidence with each sighting. Once past a threshold, one report goes to the base, or waits on board until the robot is back in range.

Animated. One ray per pixel in the box; the median hit becomes the object’s position.
Paper Fig. 19. The same pipeline on real data: detect, divide the box into pixels, cast rays into the map, filter repeated sightings.

Phones and the SubT cube were located by clustering Bluetooth signal strengths, and gas by CO₂ readings; in the paper’s example a cube was still placed within 5 m even though the robot never passed right by it. Explorers also used their zoom-and-thermal head: an algorithm picked the fewest pan-tilt angles that would cover what the fixed cameras hadn’t seen (from 1,170 possible angles). The final course was so narrow that the head mostly made single-point scans: 25 for ANYmal 1 and 156 for ANYmal 2.

One supervisor, two screens

After the Tunnel and Urban events the team noticed that operator mistakes were a big share of their errors, and that juggling several computers made it worse. For the final, the supervisor used one computer with two interfaces: a mission-control panel to command a selected robot (explore, set an area, go home, reposition to a frontier, scan, go to a waypoint), and a monitoring and reporting view with the merged map, every robot’s position and mode, and a queue of incoming detections to accept, reject or relabel before sending to DARPA.

Paper Fig. 22b. The monitoring and reporting interface: the merged map with robot and object markers, and the queue of new detections on the side.
Animated. A lower gas threshold means many more reports for the one person who has to check them all. Counts are drawn, except the real total of 30.
Scene 8

The final

In short: 23 points, the same as CSIRO Data61, and a win because CERBERUS reported its last point earlier.

The Final Event ran on 21–24 September 2021 in the Louisville Mega Cavern, Kentucky: one entrance down an inclined tunnel into an opening with three doorways to tunnel, urban and cave sections, all interconnected over several levels.

Two preliminary runs

Each team had two 30-minute practice runs with 20 objects; the course changed between runs and no maps could be carried over. On day one CERBERUS sent three ANYmals and the rover with cautious settings. The Explorer got a foot stuck in a railway track and shut down from over-current; a Carrier needed two waypoints from the supervisor because the cautious planner refused a short staircase and its LiDAR never saw the bottom of a long one. Score: 7, joint lead with CSIRO Data61. They also learned the course started with a narrow tunnel, so launching drones from the start was risky and of little use. On day two no robot was allowed out of radio range, so they explored less and scored 6 (third overall across both practice runs).

The Prize Round

On 23 September each team had one 60-minute run, with 40 objects and 45 report attempts, on a course changed again, now with a hidden mechanism that closed an opening behind a robot and objects that fell from the ceiling. Teams were kept apart so no one could learn from another’s run. CERBERUS chose the day-one settings and decided not to deploy last-minute fixes to planning and mapping, a choice they believe was crucial.

The supervisor sent the two Explorers first (ANYmal 2 into the cave, ANYmal 1 into the tunnel), parked the rover at the three doorways with its antenna aimed down the tunnel and urban openings, then sent Carrier ANYmal 4 to the urban section and ANYmal 3 to the cave. With 18 minutes left the team launched RMF-Owl to explore the high cave chamber, but it climbed into the upper staging area instead: there was no way to mark a 3-D “no-go” zone there, and an older gate-crossing feature wasn’t used. The safety pilot landed it. At the end CERBERUS had 23 points and had used all 45 attempts.

RobotRoleIn courseDistanceScoredOperating
ANYmal 1Explorer57 min240 m831.1 %
ANYmal 2Explorer60 min687 m869 %
ANYmal 3Carrier47 min311 m134.1 %
ANYmal 4Carrier50 min500 m657.5 %

What this means: the four walking robots scored all 23 points between them, covering 1.7 km in total. “Operating” is the share of time in the course a robot was actually moving or working; ANYmal 1’s is low because it spent the last half-hour trapped (see lessons).

Paper Fig. 24. The winning run minute by minute: blue is autonomous, orange is following the supervisor’s waypoints, hatched is out of radio contact; yellow stars on the top line are the 23 scored points. Note the gap between minutes 15 and 33.

The gap in that timeline is the CO₂ story from scene 7: ANYmal 4 had sent 30 CO₂ reports between minutes 10 and 15, because the team had set the gas threshold at 1500 ppm rather than DARPA’s 3000 ppm to catch leaks from further away. Each time the robot got closer, the cluster moved and another report went out, and the supervisor had to sort them all while also running the robots. ANYmal 2 later sent 11 more between minutes 37 and 42. In the last 7 minutes the supervisor stopped giving new goals and only chose which reports to submit.

Paper Fig. 27. What each ANYmal mapped on board during the run (top), and the same paths laid over DARPA’s ground-truth map afterwards (bottom).
Paper Fig. 25. The 23 scored objects on DARPA’s ground-truth map, with the robots’ camera views. Two were detected with the wrong class and relabelled by the supervisor before submitting; one was spotted by the supervisor in a video stream.
After the win

Lessons learned

In short: the hardest problems were not single algorithms but the seams between them, the human’s workload, and the gap between the lab and a real hole in the ground.

What went wrong in the final

  • Obstacles that moved. The homing and repositioning routines weren’t built for obstacles that appear after the robot has passed. ANYmal 1 met one while repositioning, started retracing its path, and stayed stuck for the rest of the run: only the supervisor could stop that retreat, and the robot was out of radio range, so he never knew. ANYmal 4 got stuck the same way but in range, and the supervisor restarted exploration.
  • Two planners that didn’t talk. The exploration planner, tuned to be optimistic, planned several times through a gap between two construction cones; the cautious navigation planner correctly refused every time. Safe, but there was no feedback, so time was lost.
  • Low ceilings and a half-second lag. To see slopes and stairs, the height map accepted points above robot height, which turned low ceilings into ghost obstacles until the robot got close. A delay of up to 500 ms between planning a path and following it made ANYmal 2 clip a scaffolding pole, which then vanished from its map while too close to see, and it was stuck for about 20 s.
  • A white corridor. The team’s camera features (BRISK) found few good landmarks in narrow, over-exposed or texture-less passages. Re-processing the run with a learned feature detector (SuperPoint) gave a clearly better multi-robot map.
  • An overloaded supervisor. Switching to another robot meant requesting its map and waiting for it to stream, several seconds each time. A single interface for control, overview and reports would have helped. Partly because of this, only 3 of the 8 breadcrumbs were dropped.
Paper Fig. 30. Where ANYmal 1 (tunnel) and ANYmal 4 (urban) ran into obstacles that appeared during the run.
Paper Fig. 33. The course’s perception traps from the robots’ own cameras: large dark spaces beyond the LEDs’ roughly 7 m reach, narrow passages, and machine-made smoke.

What the team took away

  • Legs plus wings. Quadrupeds were the main vehicle for long, large-scale exploration; drones are the only way up vertical shafts but have little endurance and payload. The team’s answer is a “marsupial” pair, a drone that rides on the walking robot, which they finished only after the competition.
  • Legged robots came of age. Walking speed rose from 0.3 m/s (2019) to 0.45 (2020) to 0.7 m/s in the final, deliberately capped although the robot could exceed 1 m/s, with no falls. The authors call it a tipping point, and list what is still missing: climbing like animals do, planning and locomotion that are less separated, a semantic sense of terrain, and using legs to touch or move things.
  • Perception must be multi-modal, and failures handled as a whole. The hand-written rules in CompSLAM’s hierarchy worked but may not generalise; the multi-robot map needs less hand-tuning and should be tested on its own.
  • Autonomy that doesn’t need the radio. Robust single-robot autonomy made communication failures bearable. Multi-robot coordination was developed but never deployed; planning with semantics (doors, stairs) and explicit risk is next.
  • Test in the field, often, and freeze before it counts. Weekly “shakeouts” (about 20-minute mock missions in a basement), at least one underground field test a month in the final year across 15 sites in four countries, a strict rule that anything failing the day-before check stays home, and no untested changes before the Prize Round.
  • More simulation. Hand-labelling detector data was slow; photo-realistic simulation with smoke, lighting and radio effects could help, but no simulator replaces field data for sensors like thermal cameras.
Animated. Two numbers from the lessons section: ANYmal’s speed at each event, and how many teams bet on legs.
Paper Fig. 23. Some of the mines, bunkers and caves in Switzerland, the USA, Norway and the UK where the team rehearsed full missions.

The team released much of its software as open source, including GBPlanner2, the ANYmal navigation planner (ART Planner), GPU elevation mapping, maplab with the M3RM server, and simulation models of its robots in the SubT virtual environment, as well as a dataset from the Prize Round.