Science Robotics · 2022 · an animated walkthrough

CERBERUS in the DARPA Subterranean Challenge

The paper in 107 seconds · narrated · sound on
Transcript

Deep underground, a four-legged robot walks out of the dark. Nobody is steering it. It maps and picks its own way, searching for survivors and lost objects. Its team won the DARPA Subterranean Challenge. How?

After a disaster, mines, caves and tunnels can be deadly for rescuers. Robots could go first. In DARPA’s final, teams had one hour to find 40 hidden objects, each within 5 metres. Underground it’s dark and smoky, radio fades, and the floor is rails, stairs and rubble.

Team CERBERUS bet on legs and wings: ANYmal walking robots, drones for shafts, and robots that drop radio relays. One planner guides them all: go where the sensors see the most unknown space, and remember skipped openings, called frontiers, for later. The walking controller, trained in simulation, reads an elevation map (a height map of the ground) and ignores it when smoke makes it lie. Position estimates hop from sensor to sensor, skipping any that fails, so a robot never loses track. Cameras spot survivors and backpacks; phones give themselves away by Bluetooth.

In the final, four ANYmals walked over 1.7 km and found 23 of 40 objects. Tied with CSIRO Data61, CERBERUS won by reporting its last find earlier. Here are the real run’s last minutes, sped up: the score ticks up to 23. The lesson: legs plus wings, backup sensors, and supervised autonomy: robots decide each step, one human sets the strategy.

Footage: Robotic Systems Lab, ETH Zurich (“ANYmal at DARPA SubT final run”, “Team CERBERUS Wins the DARPA Subterranean Challenge”) and Team CERBERUS (“CERBERUS: DARPA Subterranean Challenge Finals Prize Run in 2min”), which show the DARPA Final Event broadcast view and the team’s own robot view. Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video. Animations are schematic.

The story in plain words

Four-legged robots from Team CERBERUS explored an unknown underground course for an hour, largely on their own, found 23 of 40 hidden objects and won the DARPA Subterranean Challenge. This paper explains the ideas behind that run.

  1. Why this matters

    After a collapse, a fire or a flood, mines, caves, tunnels and underground stations are dangerous places for rescuers. Robots could go in first and tell a small team what is down there: where the passages lead, where people or hazards are.

  2. What makes it hard

    Underground it is dark, dusty and sometimes smoky; one tunnel looks like the next; radio signals die behind rock; floors have rails, stairs, rubble and mud. Imagine exploring an unknown building in the dark with a walkie-talkie that stops working after a few rooms, while only one person outside may help.

  3. What people did before

    DARPA ran a three-year, $82 million contest to push this forward, with 20 teams from 11 countries. Teams tried wheels, tracks, legs and drones. Mapping and planning methods had come a long way, but underground they still had “many reasons to break”.

  4. What this paper does

    It lays out the recipe of Team CERBERUS: four-legged ANYmal robots as the main explorers, drones for shafts, robots that drop their own radio relays, one exploration planner shared by all robots, a learned walking controller, position tracking that survives failed sensors, and a single human who sets goals while the robots make every step’s decision.

  5. What they showed

    In the Final Event (Louisville Mega Cavern, September 2021), four ANYmal robots walked more than 1.7 km in 60 minutes and correctly reported 23 of 40 objects. CSIRO Data61 also scored 23; CERBERUS won the tie-break because its last artifact report came with more than a minute left, CSIRO’s with under 30 seconds.

  6. Why it’s a step forward

    The lessons: legs plus wings suit the underground; resilience comes from sensors that back each other up; and a human works best as a strategist, not a driver. The authors expect the same recipe to help at construction sites and industrial inspections. Still missing: robots that understand what they see, weigh risks, and coordinate without a central base.

Words used below
artifact
DARPA’s word for a hidden object to find (mannequin “survivor”, phone, backpack…)
frontier
the edge between mapped and unmapped space: an opening not yet explored
elevation map
a grid of ground heights around the robot, used to plan footsteps
SLAM
building a map while working out where you are in it
supervised autonomy
robots make moment-to-moment decisions; one human sets goals
1 / 8
walking robots radio link exploration paths & maps
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough. This paper is a short overview: each technology it names has its own detailed paper, cited in each section.

Scene 1

Why send robots first

In short: underground spaces are dangerous and unknown, so DARPA asked teams to build robots that could scout them for first responders, and turned it into a timed treasure hunt.

Underground settings range from tunnels several kilometres long to multi-level urban structures, from narrow cave passages and vertical shafts to vast chambers. They can be hard and dangerous for people to enter, which is why robots that give a small team of operators rapid situational awareness before anyone goes in are valuable for both military and civilian first responders.

The DARPA Subterranean (SubT) Challenge was a three-year, $82 million competition that brought together more than 300 competitors from 20 teams in 11 countries. It had a Systems track (real robots in real underground courses) and a Virtual track (software only, in simulation). The Systems track visited a mine tunnel network (Tunnel Circuit, August 2019) and multi-level urban underground structures (Urban Circuit, February 2020); the Cave Circuit (August 2020) was cancelled. The Final Event in September 2021 combined all three.

The rules that shape everything

  • Each team had 60 minutes to find as many as possible of 40 artifacts: mannequin “survivors”, cellphones, backpacks, drills, fire extinguishers, vents, gas-filled rooms, helmets, ropes and a custom “SubT cube”.
  • A report scored a point only if it named the right type and placed it within 5 m of its true position. The number of reports was limited, to discourage guessing.
  • Only one team member, the Human Supervisor, was allowed to manage and talk to the deployed robots.
Animated. Watch the third report land outside the 5 m circle: it uses up a report but earns nothing. Positions illustrative.

With one human and a radio link that fades with depth, most of the work has to be done by the robots themselves. That is the central demand of the contest, and of this paper.

Scene 2

What makes it hard

In short: DARPA named four areas where robots break down below ground, and a winning team needs all four at once.

  • Autonomy. Map, navigate and search complex, changing places “without substantial human interventions”.
  • Perception. Keep working in dust, fog, mist, water, smoke and darkness, with obscured views.
  • Networking. Get information back to the Human Supervisor despite limited line of sight, varying rock and the way radio waves behave underground.
  • Mobility. Cross narrow passages, sharp turns, big drops and climbs, slopes, steps, falling debris, mud, sand and water.
Paper Fig. 2B–E. What the robots saw during the winning run: a narrow cave passage, rough cave ground, a mine tunnel filled with smoke, and an urban corridor so uniform that cameras find little texture to track.

By the start of the challenge, teams were trying every kind of platform: wheeled, tracked, legged and flying robots. Decades of work on mapping and path planning existed, but the authors put it bluntly: underground environments “proved to be the domain where autonomy had many reasons to break”.

Scene 3

The team: legs, wings and a self-built network

In short: CERBERUS was a “system of systems”: walking robots that explore, flying robots that reach where legs can’t, and a radio network the robots lay down themselves, all run by one human.

The team was an international partnership: the University of Nevada, Reno, ETH Zurich, NTNU, UC Berkeley, the University of Oxford, Flyability and Sierra Nevada Corporation.

Paper Fig. 1. The ANYmal C SubT robot in an urban setting (A), a rail tunnel (B) and a cave (C); the tethered rover acting as a radio hub (D); a medium-sized aerial scout (E) and the caged RMF-Owl drone (F) in mines.
RobotJobKey facts from the paper
ANYmal C “SubT” ExplorerGo deep, find artifactsIP67 water/dust-proof; 55 kg with payload; 80 min continuous walking; 3 LiDARs; Alphasense Core (3 mono + 4 colour cameras); pan-tilt head with zoom camera, thermal camera, microphone and spotlight
ANYmal C “SubT” CarrierExtend the radio networkFewer sensors (1 LiDAR + Alphasense Core); carries 4 radio modules and drops them
Aerial scoutsFast, vertical explorationMedium multirotors (DJI M100 frame; LiDAR, colour and thermal cameras); caged, collision-tolerant Gagarin (built on Flyability’s Elios) and RMF-Owl; Kolibri tricopter (over 20 min flight; can pitch independently of its motion)
Tethered roverRadio hub300 m optical fibre that unrolls from the base station as it drives; high-gain directional antenna

What this means: the legged robots did the heavy exploring; everything else existed to give them reach (drones) and a voice back to the base (rover and relays).

A radio network the robots build themselves

Underground, radio propagation is hard to predict and depends heavily on the tunnel geometry. CERBERUS used an ad-hoc wireless mesh network at 5.8 GHz, in which the legged robots and the rover were themselves nodes. Each Carrier ANYmal ferried four communication-extender modules on its back and released them on the supervisor’s command by lowering its torso and switching an electromagnetic release. Each module was a sealed, water- and dust-proof unit with a Rajant BreadCrumb DX2 radio and a battery for two hours. Drones joined the mesh as clients.

Animated. The body dips, the module stays behind and starts relaying; the robot walks on with one module fewer. Motion illustrative.

Who is in charge? Supervised autonomy

The team ran in supervised autonomy: the Human Supervisor read the maps arriving at the base station and gave high-level tasks; each robot carried them out on its own. The paper lists four ways to task a robot:

  1. explore an unknown area while staying in radio range, backtracking to a connected spot if the link is lost;
  2. explore beyond radio range for a set time budget, then come back into range;
  3. walk to a chosen frontier near unmapped space, then start exploring on its own;
  4. as a last resort, go to a waypoint in the map along a path the robot finds itself.
Animated. The four tasking modes in turn; watch the signal bars over the robot as it leaves and re-enters radio range. Layout illustrative.
Scene 4

Where to go next

In short: a single exploration planner, shared by walking and flying robots, looks for the most informative path nearby and falls back on a mission-wide memory of unexplored openings.

At the core of CERBERUS’s autonomy was the graph-based exploration path planner, GBPlanner (Dang et al. 2020; Kulkarni et al. 2022). It works on a volumetric map of the environment, built from the robot’s sensor data and poses (Voxblox), and has two stages:

  • Local exploration. Given the robot’s position, the current map and bounds on the exploration space, it samples a random graph inside a local box around the robot, then picks the path along which its range sensors would see the most unknown space. The planning is truly 3-D and, for ground robots, respects limits on slopes and other terrain features.
  • Global mode. In complex, steep underground layouts the local stage can run out of useful paths. The planner then switches, automatically or when the supervisor asks, to a sparse graph built up over the whole known map. It stores the frontiers of exploration and the paths to them, so it can reposition the robot to unexplored areas, or return it home or to the last point with a radio link when its battery or time budget runs out.
Animated. Going back reveals nothing; the side opening a little; ahead the most, so that becomes the next path. Numbers illustrative.
Paper Fig. 2F–G. Left: a robot’s exploration path (yellow) through a cave. Right: that path adjusted into a traversable path (blue, with axis markers) on the elevation map; white cells are steppable.
Real run. ANYmal 2 pushing into the cave section of the final, from the team’s replay at 20× speed: its four camera views on top, and below, the 3-D map it builds as it goes. Footage: “Team CERBERUS Wins the DARPA Subterranean Challenge”, Robotic Systems Lab, ETH Zurich.

Because the same planner drives legged and flying robots, the team could add or swap platforms without writing a new exploration strategy for each, which the authors list as one of their lessons.

Scene 5

Walking it

In short: a legged-robot planner reshapes the exploration path to fit the ground, and a walking controller learned in simulation executes it, deciding for itself how much to trust its view of the terrain.

The final course had rails and hoses in the tunnel section, stairs in the urban section and a rough, slippery cave with gravel, rocks and a steep, narrow passage. Two layers turned an exploration path into safe footsteps:

  • Navigation planner. A reachability-based planner for legged robots (Wellhausen & Hutter 2021) follows the exploration waypoints but can adjust the path locally. It uses an elevation map computed on the GPU (Miki et al. 2022) and a cost map of how traversable each patch is.
  • Learned walking controller. A perceptive locomotion policy (Miki et al., Science Robotics 2022) tracks the velocity commands. It was trained entirely in simulation with reinforcement learning. To make it robust, training randomised external forces and torques on the body, the initial body and joint velocities, the total mass and the slipperiness of the ground.

The key property is how it combines its senses. The elevation map is essential for obstacles such as stairs, but dust, smoke and puddles add noise to the depth sensors. The controller was trained with a variety of noise in its terrain input and learned, end to end, how much to rely on each source: use the map when it looks reliable, for fast and smooth walking; ignore it when it looks degraded. The same training lets it cope when a sensor stops sending data altogether.

Animated. In the second half, red map points float above the steps (fake bumps from smoke) and the gauge drops, while the robot keeps climbing. Gauge illustrative.
Real run. An ANYmal stepping over the uneven, gravelly cave floor during the final (a close-up crop of the DARPA broadcast view). Footage: “ANYmal at DARPA SubT final run”, Robotic Systems Lab, ETH Zurich.

Earlier in the competition the team also tried different feet for ANYmal: point feet, flat feet and wheels. They settled on point feet as the most versatile across terrain types, avoiding the need for specialised designs. The drones used conventional cascaded position–attitude control, with fixed gains or model predictive control.

Scene 6

Where am I?

In short: every robot tracks its own position with a chain of sensor-specific estimators that skips any link that fails, and the base station then fuses all robots into one mission map.

Localisation and mapping (SLAM) rested on two ideas: sensors of different kinds fail in different situations, so combining them gives redundancy; and the software had to adapt to very different robots, from a 55 kg quadruped to a small drone. The result was a two-stage solution.

On each robot: CompSLAM

CompSLAM (Khattak et al. 2020) fuses visual and thermal images, LiDAR depth, inertial measurements and the legs’ kinematic estimates in a hierarchy. Each sensing modality (or group) has its own pose estimator. If its output passes data- and process-level health checks, it becomes the starting point for the next estimator in the hierarchy. If an estimator is judged inconsistent, the previous estimate is handed straight to the next one, skipping the failed step. So every robot always has an up-to-date pose and map.

Animated. After four seconds the camera estimator fails its check; the estimate flows around it and the output never stops. Order of the chain simplified.

At the base: M3RM, one map for the team

When in radio range, each robot incrementally sends its odometry and sensor data to the base station: visual-inertial factor graphs plus subsampled, compressed LiDAR scans, to fit the network’s limited bandwidth. The base station, with far more computing power, optimises them jointly into a globally consistent multi-robot map (M3RM, built on the maplab framework) and uses it to report optimised artifact positions. If that fails, the per-robot CompSLAM maps serve as a fallback; most drones reported only their onboard estimates. Splitting the job between robot and mission level avoids a single point of failure.

Scene 7

Spotting objects

In short: scoring means naming the object and placing it within 5 m, so detection was paired with careful localisation, and a human confirmed every report.

All visual artifacts (everything except cellphones, gas and the SubT cube) were detected with YOLOv3, a fast image-detection neural network, trained on more than 40,000 labelled images collected underground. The team deliberately gathered images with obscurants, from different cameras, against varied backgrounds, under changing light and with motion blur. The network ran onboard every robot.

For each detection, the robot cast rays from the pixels inside the detection box into its volumetric map to get a first 3-D position, then refined it with later sightings as it moved. A binary probability filter, updated with every detection, suppressed false alarms.

Animated. Early rays scatter; later ones land closer together, and the estimate settles on the object. Spread illustrative.

Objects that cannot be seen were sensed differently: the cellphone and SubT cube by their Bluetooth signal strength, and gas with a CO₂ sensor. Within a window of measurements, the object was placed at the robot’s position where the signal was strongest. In every case the Human Supervisor inspected the detection and accepted or rejected it before a report was sent.

Scene 8

The winning run

In short: CERBERUS tied for the most points and won on the tie-break, because its final report came earlier.

The Prize Round of the Final Event took place in the Louisville Mega Cavern, Kentucky, in September 2021, with eight Systems-track teams and nine Virtual-track teams. In the winning run CERBERUS deployed four ANYmal C SubT robots. Together they walked more than 1.7 km and correctly reported 23 of the 40 artifacts.

Systems-track teamRobots (as described in the paper)Points
CERBERUSlegged, flying, tethered rover23
CSIRO Data61tracked, legged, drones23
MARBLElegged, wheeled18
Explorermainly custom wheeled and flying17
CoSTARlegged, wheeled, drones13
CTU-CRAS-NORLABlegged, tracked, flying7
Coordinated Roboticswheeled, flying2
Robotikawheeled2

What this means: the top two teams were level, and the next best found 5 fewer artifacts. The tie-break rule, “the team that identified its last artifact earliest would win”, decided it: CSIRO Data61’s final point came with under 30 seconds left, CERBERUS’s with more than a minute left. CERBERUS won the challenge and the $2 million prize.

Paper Fig. 2A. The team map from the winning run: each ANYmal’s map in its own colour over DARPA’s ground-truth map (yellow), with the 23 correctly reported artifacts labelled. Insets show a detected survivor and backpack.
Real run. The last minutes of the prize run in the DARPA broadcast view, sped up: the score at left ticks from 21 to 22 to 23, and the clock runs out with the team at the base station. Footage: “CERBERUS: DARPA Subterranean Challenge Finals Prize Run in 2min”, Team CERBERUS.
Wrap-up

Lessons and limits

In short: legs plus wings, redundant perception and a human strategist worked, but robots still lack understanding, risk awareness and decentralised teamwork.

  • Legged and flying robots are the right mix underground. CERBERUS relied almost entirely on walking and flying systems from the start, and by the Final almost all teams leaned heavily on them. Walking robots were the main explorers, with endurance and terrain ability beyond wheels or tracks; drones filled the gaps: vertical spaces and impassable ground. But small drones have short endurance, and the places that need them may be far from the entrance. The team worked on ANYmal carrying the RMF-Owl drone (“marsupial” deployment), finished only after the competition.
  • Resilient perception and autonomy are the key technologies. Fusing sensors that fail in different ways (LiDAR does not need visual texture; cameras do) minimised breaking points. And one exploration policy that reflects the geometry of underground spaces could guide very different robots.
  • Autonomy and human input work best together. Without prior infrastructure, underground spaces don’t allow reliable high-bandwidth links, so autonomy is what makes missions possible. When a link exists, the supervisor’s high-level commands improve where robots go. All the dense, moment-to-moment path decisions stayed with the robots.

What is still missing (the authors’ list)

  • semantic understanding of the environment;
  • comprehensive modelling of navigation risks, for example how perception conditions should influence path choices;
  • tighter coupling between exploration planning and traversability;
  • multi-robot SLAM that stays consistent under severe sensor degradation;
  • efficient decentralised coordination of many robots.

The authors expect the same methods to transfer beyond the underground, to mapping construction sites, exploring new infrastructure, and inspecting ships and process tanks: places with the same mix of hard terrain, poor sensing and a need for autonomous mapping and object recognition.

About this paper: it is a short overview (a few pages, two figures, no experiments beyond the competition result). Numbers on this page are the ones it reports; the detailed evaluations live in the cited papers on GBPlanner, CompSLAM, M3RM, the navigation planner, elevation mapping and perceptive locomotion, and in the team’s Field Robotics paper on the Tunnel and Urban circuits.