RSS 2024 · Robotics: Science and Systems · an animated walkthrough

Rethinking Robustness Assessment: Adversarial Attacks on Learning-based Quadrupedal Locomotion Controllers

The paper in 102 seconds · narrated · sound on
Transcript

This controller helped a robot team win a competition in real caves. Yet a learned attacker tripped it on flat ground, with tilt errors of just three degrees. On a wet board, the original slips and falls. Trained against its attacker, it holds. So where are a robot’s hidden weak spots?

Four-legged robots are meant for dull, dirty and dangerous work. Their legs are run by a neural network trained in simulation, and nobody can read off when it will fail. Standard push tests rarely find the weak spot: in the paper, zero falls in a thousand trials. A learned attacker with the same force limit made it fall every time.

The idea: a second AI as a sparring partner. By trial and error it learns to nudge commands, bend the robot’s sense of its own tilt and push its feet, within realistic limits. Its moves are kept smooth, like real sensor drift.

Then they attacked a champion, the controller used in the DARPA Subterranean Challenge. With tilt errors up to two degrees, it held. Three degrees toppled it, in simulation and on the real robot. Of a hundred roboticists with joysticks, only three succeeded, and the attacker was faster and hit harder.

The fix: keep training the controller against its attacker. New attackers then failed even at five degrees, on the real robot too, and it walked just as well. It still crossed rough ground outdoors. Instead of hoping a robot is robust, we can search for its worst case, and patch it.

Footage: Robotic Systems Lab, ETH Zurich (the paper’s video, YouTube); the cave clip is DARPA Subterranean Challenge footage shown in that video. Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video. Animations are schematic.

The story in plain words

The authors trained a second AI to make walking robots fall: it found gentle, realistic moves that topple even a DARPA-winning controller, and training against it made that controller tougher.

  1. Why this matters

    Four-legged robots are meant for dull, dirty and dangerous jobs: inspecting tunnels, carrying loads, searching collapsed buildings. Their leg controllers are now mostly neural networks trained in simulation, and one of them helped ANYmal robots win the DARPA Subterranean Challenge in real caves. Before deploying many such robots, we need to know when they fail.

  2. What makes it hard

    A neural network is a black box: no formula says when it breaks. Real failures are rare and come from sequences of small things: a tilt reading that drifts, a command at the wrong moment, a slip. It is like finding the combination of a lock with many dials that must be turned in rhythm, 50 times a second.

  3. What people did before

    Push the robot with fixed or random forces; randomize the simulator during training so the controller copes on average; let engineers kick and pull it. These catch big shoves but don’t search for the subtle worst case. Earlier learned attacks targeted self-driving cars, or a simplified simulated quadruped without tests on a real robot.

  4. What this paper does

    It trains a second AI, an adversary, by trial and error to make the robot fall. Like a sparring partner that is not allowed to cheat, it may only do realistic things: bend the robot’s own tilt estimate by at most 3°, send velocity commands within ±0.5 m/s, and push the feet with up to 15 N in simulation. Its moves are kept smooth, like real sensor drift.

  5. What they showed

    Four standard push tests: 0 falls in 1000 trials. The adversary: a fall every time. The DARPA controller fell on flat ground with tilt errors of 3°, about what its tilt estimate already gets wrong on a plain floor, both in simulation and on the real ANYmal. Of 100 roboticists with joysticks, 3 managed a fall; the adversary was about 3 times faster than the quickest of them.

  6. Why it’s a step forward

    Robustness becomes something you search for rather than hope for. Training against the adversary produced a controller that new attackers could not topple even at 5°, on hardware too, while it walked just as well. Limits: only the robot’s body sensors were attacked (not its terrain map), and no method can promise to find every weakness.

Words used below
Controller (policy)
the neural network that turns sensor readings into leg motions
Reinforcement learning
learning by trial and error, guided by a score
Adversary
a second AI trained to make the controller fail
State estimator
software that works out the robot’s tilt from its motion sensor and legs
Domain randomization
randomly varying pushes, friction and noise in simulation during training
Fine-tuning
continuing to train an existing controller on new situations
1 / 7
walking controller standard tests · humans learned adversary commands tilt errors pushes
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough.

Scene 1

Why it matters

In short: legged robots are heading into places where a fall is costly, and their brains are neural networks nobody can inspect line by line, so we need a systematic way to find where they break.

Quadruped robots are expected to take over “3D” work, meaning dull, dirty and dangerous: transport, underground inspection, rescue. What makes that possible today is the controller, the program that decides every 20 milliseconds how to move twelve joints. Modern controllers are neural networks trained with reinforcement learning in huge parallel simulations. Instead of an engineer writing down the robot’s physics and solving it online, the network learns a mapping from sensor readings to motor commands by trial and error.

These controllers are remarkably tough. One of them, the perceptive controller of Miki et al. (2022), which this paper calls the Miki policy, walked ANYmal robots through caves, tunnels and snowy mountains and was part of team CERBERUS’s win at the DARPA Subterranean Challenge. Failures in the field are rare.

Rare is not never. Neural networks are known to be sensitive to small, well-chosen changes in their input (the classic example is an image classifier fooled by invisible pixel noise). For a walking robot the question becomes: is there a sequence of small, realistic disturbances that makes this controller fall? The paper argues that the answer matters for safety, and that the mildest such sequence is itself a useful measure of how robust a controller really is.

Paper Fig. 1. The DARPA-winning controller on a plain, flat lab floor, tipping over under a sequence of small attacks on its tilt estimate and velocity command. Nothing pushes it physically.
Scene 2

Hidden weak spots

In short: standard push tests check whether a robot survives big shoves, but a failure usually needs the right small things at the right moments, and only a search finds that.

Why is finding the weak spot hard? The space of possible disturbances is high-dimensional (several commands, three tilt angles, forces on four feet) and sequential (a new value every 0.02 s). Trying random combinations is hopeless, the “curse of dimensionality”. And a controller trained to be robust has already learned to shrug off the random noise it saw in training.

The existing toolbox, which the paper treats fairly, consists of:

  • Standard push-recovery tests: apply fixed or random forces to the robot’s body and count falls. Useful to compare controllers against large pushes, but they cover a narrow set of scenarios.
  • Domain randomization: randomize masses, friction, noise and pushes in simulation so the controller works “on average”. The paper later shows this is not enough (section 8).
  • Search-based testing, well established for self-driving cars, where reinforcement learning finds the pedestrian movements that make a car crash. For quadrupeds, a prior study used evolutionary search on joint torques, but on a simplified simulated “Ant” model and without a real robot.

To show the gap, the paper uses a deliberately simple “didactic” controller (the open-source blind controller of Rudin et al., a small network that sees only the robot’s own body sensors), walking forward at 0.4 m/s in simulation. Four standard tests with the same force limits as the adversary were run 1000 times each:

Test (pushes on the body, same range)Falls
ST 1 · a random constant push0 / 1000
ST 2 · random pushes, re-drawn every 2 s (0.5 Hz)0 / 1000
ST 3 · maximum-strength pushes in random directions, every 2 s0 / 1000
ST 4 · a random impact, held for 0.2 s0 / 1000
Learned adversary100 %

What this means: a controller can pass every standard test and still have a weakness that a searching attacker finds every single time.

Animated, illustrative. Why timing matters, in a toy balance model: the same maximum push, applied at random (left) or in rhythm with the robot’s sway (right). This is an intuition for sequential attacks, not the paper’s analysis of what its adversary does.

The paper also runs an attack–patch–re-attack loop on this controller. First the adversary may only push from one side and finds a fall; the controller is then fine-tuned against it; a new adversary allowed both sides immediately finds the side that was never patched. Section 7 shows a cleaner version of this experiment.

Scene 3

The sparring partner

In short: the attacker is itself a small neural network, trained by trial and error to make the robot fall, but only with disturbances a real robot could plausibly meet.

The key design decision is what the attacker may touch. The authors picked three sources of failure from field experience, each with a realistic range and a limit on how fast it may change (at most 10 % of its range per 0.02 s step):

ChannelWhat it stands forRange (Miki policy)
Commandsa careless operator or navigation module sending velocity commandsx, y: ±0.5 m/s · yaw: ±0.5 rad/s
Observationserrors in the robot’s own tilt estimate (roll, pitch, yaw) from its motion sensor and leg kinematics. Joint encoders are precise, so they are left alone.±3.0°
Perturbationsunseen forces: pushes and payloads on the body, stumbles and slips at the feet. Only possible in simulation.feet: ±15 N

What this means: no attack exceeds what a working robot routinely experiences; the paper rules out things like a 10 000 N shove on a 50 kg robot, which would teach nothing.

The adversary sees the same inputs as the controller it attacks and outputs its disturbances 50 times a second. It is trained with PPO, a standard reinforcement-learning algorithm. The walking controller is never changed during the attack and needs no special access: it is treated as a black box.

What the attacker is rewarded for

The attacker’s score has one main term and three helpers:

score per step = −(1 if the robot is still up) + bonus(tilted body) + bonus(shaking body) + bonus(motor torque above its limit)

The first term makes a fall the goal: every step the robot survives costs the attacker. The bonuses give a signal before the first fall happens, which makes learning much faster. On top of this sits the smoothness penalty of section 4.

Why not simply use the negative of the walking controller’s own training reward, the textbook “zero-sum” game? Because that reward is dominated by performance terms such as how well speed commands are tracked, foot clearance and energy. An attacker minimizing it learns to make the robot walk badly, not to make it fall. A fall even ends the episode, which in that reward is worth less than a long stretch of sloppy walking. The ablation confirms this:

Animated, numbers from Table IV. Only scores that include the “still standing” penalty find a fall. Dropping the helper bonuses still works but needs about 3× more training.
Attacker’s scoreIterations to convergeRobot falls?
Proposed (fall penalty + helpers)240yes
Zero-sum: negative walking reward320no
Without the “still standing” penalty280no
Without tilt and shake bonuses650yes
Without the torque bonus720yes

What this means: for safety testing, the attacker must be rewarded for the failure you care about, not for making the robot’s task harder in general.

Why three channels at once

Earlier legged-robot studies attacked one channel at a time. With the small ranges above, the paper finds that no single channel is enough to topple the Miki policy, but every combination of two or three is (Table III). Real failures of a mature controller also tend to come from several factors together.

Animated; results from Table III (robot motion illustrative). Commands, tilt errors or foot pushes alone: the controller holds. Any pair, or all three: it falls.
Scene 4

Smooth attacks

In short: an unconstrained attacker learns to flicker its signal between the extremes; a smoothness penalty makes it produce slow drifts that look like real sensor errors.

Reinforcement learning happily finds “bang-bang” solutions: the attacker slams the tilt error from +3° to −3° and back many times a second, which is effective in simulation but tells us little, because no real estimator behaves that way. The authors measured their robot’s tilt estimate against motion capture while it walked on a flat indoor floor: the error drifts slowly and peaks at almost 3°, and it grows on slippery or uneven ground or under pushes.

The fix is Lipschitz regularization. A network’s Lipschitz constant bounds how much its output can change when its input changes a little. For a stack of layers with ReLU-type activations, the product of each layer’s weight norms (here, infinity norms) is an upper bound on that constant. Adding this product, times a weight λ, as a penalty to the attacker’s training loss keeps the constant small, so the attack becomes smooth.

maximize   E[ Σt γt (−1alive + rtaux) ] − λ · Πi ‖θi‖∞  // fall-seeking score, minus a cap on how sharply the attacker can react
Paper Fig. 2. Two learned tilt attacks that both topple the robot in simulation. Top: plain reinforcement learning, flipping between ±3° all the time. Bottom: with Lipschitz regularization, slow ramps and plateaus.
Paper Fig. 5. What real errors look like: the robot’s estimated roll (blue) against motion-capture ground truth (grey dashed) on flat indoor ground. The gap wanders and reaches almost 3°, the same size as the attack budget.
Scene 5

Toppling the champion

In short: with only command nudges and tilt errors of 3°, the attacker makes the DARPA-winning controller fall on flat ground, in simulation and on the real robot alike.

The second case study attacks the Miki policy, a perceptive controller with a recurrent memory that sees a 208-point height map of the terrain around the robot and is known from field use to be robust to sensor noise and pushes. The attacker (a network with three hidden layers) sees the same inputs.

For the real-robot test, only the two channels that can be reproduced on hardware were used: velocity commands and tilt-estimate errors. The same learned attacker was then run against the real ANYmal walking on a flat lab floor, with a safety stop that caught the robot before it hit the ground. To find the smallest attack that works, the tilt error was rescaled to different maximum sizes:

Maximum tilt error0°1°2°3°5°
Simulationholdsholdsholdsfallsfalls
Real ANYmalholdsholdsholdsfallsfalls

What this means: the weakness found in simulation is real: the threshold is the same on hardware (between 2° and 3°), and it sits inside the range of errors the estimator produces anyway.

Real run. The real ANYmal on the flat lab floor while the learned attacker nudges its velocity commands and bends its tilt estimate by at most 3°: within about five seconds its front legs buckle and it sinks toward the floor. Footage: Robotic Systems Lab, ETH Zurich (from the paper’s video).
Paper Fig. 6. Simulation versus real robot under the same attack. Top: the table above (“Fail/Succeed” refers to the attack). Below: the robot falling at 3° in simulation and on hardware, with similar joint trajectories.
Paper Fig. 4. The attack that toppled the real robot: velocity commands (top) and tilt-estimate errors within 3° (bottom), over about three seconds.
Scene 6

Humans vs machine

In short: experienced roboticists poking the robot by hand rarely find the weakness; the learned attacker finds it faster and exploits it harder.

A common way to harden a controller is intuition: an engineer kicks the robot, adds noise, watches. Is that enough? The authors organized an AI safety challenge. In simulation, the simple didactic controller walked forward at 0.4 m/s. Participants used joysticks to push its body (up to 100 N along each horizontal axis) and to override its sideways and turning commands (within 0.5 m/s and 0.5 rad/s). For a fair match, the machine attacker got no joint readings, since humans only see the rendered robot, and could act only 5 times per second.

100 volunteers with a robotics background each had 5 minutes. Three made the robot fall.

Winner 1Winner 2Winner 3Adversary
Average survival time4.0 s5.8 s11.5 s1.38 s
Average roll rate while falling1.77 rad/s1.53 rad/s2.80 rad/s5.04 rad/s

What this means: the adversary topples the robot about three times faster than the best human and with a much more violent fall, and 97 of 100 experts found nothing at all.

Paper Fig. 11. The challenge setup: a joystick drives pushes and commands in the simulator (left); snapshots of one of the three successful attempts (right).
Scene 7

Train against it

In short: the attacks found in testing become training material; the patched controller resists new attacks, on hardware too, and walks as well as before.

Closing the loop is simple. Training of the controller resumes from its last version (actor and critic networks warm-started). Most simulated robots keep the usual random disturbances, and some meet the learned adversary instead. The objective does not change; only the situations the controller practices on do. Too many adversarial robots make the controller over-cautious, so the share is kept small.

Animated, illustrative. For the Miki policy, each simulated robot had a 5 % chance of meeting the adversary (trained on all three channels) instead of random perturbations.

The patched Miki policy

New attackers were then trained against the fine-tuned controller, with a larger budget of up to 5° tilt error. None could topple it, neither in simulation nor on the real robot, even while an operator pushed it by hand in addition (Fig. 9). On a wet whiteboard and a soft mattress, where the original controller quickly lost balance under perturbation, the fine-tuned one reacted and recovered (Fig. 7). Outdoors it still crossed rough slopes, slippery ditches, sand and steep gravel roads (Fig. 8).

Paper Fig. 9. Re-attacking the fine-tuned controller: attacks from 0.1° to 5° all fail, in simulation and on the real robot, with an operator adding pushes.
Paper Fig. 7. Indoor stress test on a wet board and a mattress. Top: original controller. Bottom: after adversarial fine-tuning.
Real run. The fine-tuned controller on and around a soft mattress while two people pull on it with ropes: it keeps its footing. Footage: Robotic Systems Lab, ETH Zurich (from the paper’s video).
Paper Fig. 8. The fine-tuned controller outdoors: robustness did not cost it its ability to cross hard terrain.
Real run. The fine-tuned controller crossing a rough, grassy slope into a wooded path, one of the outdoor tests. Footage: Robotic Systems Lab, ETH Zurich (from the paper’s video).

Did robustness cost performance? Compared with the original controller in simulation:

TerrainChange in avg. joint torqueChange in avg. tracking error
Flat+0.42 Nm+0.002 m/s
Rough−1.76 Nm−0.006 m/s
Standard stairs+0.26 Nm+0.008 m/s
Random stones−0.26 Nm+0.006 m/s

What this means: speed tracking changes by less than a centimetre per second and effort barely moves, so the extra robustness came essentially for free.

Patching one weakness at a time

A controlled experiment shows the mechanism. The didactic controller was trained with random pushes that only came from within 45° of straight ahead or behind, leaving its sides “unguarded”. Push tests confirmed it was weak sideways. The first adversary found a fall by pushing it leftward; after fine-tuning against it, the left was patched but the right was still weak. A second adversary found exactly that, and after a second round of fine-tuning the controller resisted pushes from both sides.

Paper Fig. 14. Two rounds of attack and fine-tuning. The polar plot (b) shows the smallest constant push that causes a fall, by direction: initial (grey), after round 1 (orange), after round 2 (blue). Snapshots show the leftward (c) and rightward (d) attacks.
Analyses

More findings

In short: randomizing the simulator is not enough, the smallest successful attack can serve as a robustness score, attackers that can see terrain use it, attacks can be steered toward different failures, and the method also breaks a classical model-based controller.

Is domain randomization enough? No.

A controller trained with random pushes of up to 100 N, re-drawn 5 times a second, still fell to a learned attacker with the same 100 N limit. It survived larger pushes than a controller trained without randomization, so randomization helps, but attacks are a complement, not a replacement.

Attacks as a robustness score

Give the attacker a variable budget and shrink it until the controller survives: that smallest budget is a number you can compare across controllers and terrains. Three controllers were compared: no randomization (“policy”), randomized pushes (“policy-DR”), and randomized pushes plus adversarial fine-tuning (“policy-DR-ft”). The paper reports the ordering policy < policy-DR < policy-DR-ft on every terrain, with the biggest gains from fine-tuning on uneven terrain. The exact values are only given in the chart.

Paper Fig. 10. Smallest push range that topples each controller, per terrain (larger = sturdier).

Attackers that use the terrain

When the attacker also sees the height map, it exploits the ground: on stairs it first turns the robot toward the edge, then attacks. On flat ground it attacks directly. The real-robot attacks were done on flat ground for safety, so uneven terrain may hide even easier failures.

Paper Fig. 12. A terrain-aware attacker on stairs (a) versus flat ground (b).

Different failures on demand

One training run converges to one attack strategy, but a controller may have several weaknesses. Changing the weights of the attacker’s bonus terms steered it to different kinds of damage: falling over, the body hitting a front leg (self-collision), or a knee motor exceeding its torque limit during a single-foot touchdown. (The paper’s image of these three failures, Fig. 13, is omitted here because it did not convert cleanly.)

Not only neural controllers

The method only needs to run the controller, not understand it. The authors attacked an open-source model predictive controller (Kim et al.’s MIT Mini Cheetah controller, adapted to the Unitree A1) through commands and observations with the same ranges and score, and toppled it in simulation.

Paper Fig. 15. A learned attack toppling a model predictive controller on the Unitree A1 in simulation.
Honest limits

What it doesn’t do yet

In short: this is a first step toward measuring robustness, not a certificate.

  • Only body sensors were attacked. The Miki policy also reads a 208-number height map of the terrain, which is probably vulnerable too; generating such high-dimensional attack sequences may need generative models rather than reinforcement learning.
  • No guarantee of completeness. A learned attacker finds a weakness, not all of them. Diverse rewards help, but covering every weakness needs further work.
  • Fair comparisons are hard. Using attacks to rank controllers or training methods requires each controller to be the best its method can produce, and robots and sensors differ. The authors leave rigorous comparisons to future work.
  • Pushes stay in simulation. External forces could not be reproduced precisely on hardware, so the real-robot attacks used commands and tilt errors only.

The authors’ hope is practical: that legged-robot researchers treat worst-case testing as a standard step before large-scale deployment of neural controllers.