The authors trained a second AI to make walking robots fall: it found gentle, realistic moves that topple even a DARPA-winning controller, and training against it made that controller tougher.
Why this matters
Four-legged robots are meant for dull, dirty and dangerous jobs: inspecting tunnels, carrying loads, searching collapsed buildings. Their leg controllers are now mostly neural networks trained in simulation, and one of them helped ANYmal robots win the DARPA Subterranean Challenge in real caves. Before deploying many such robots, we need to know when they fail.
What makes it hard
A neural network is a black box: no formula says when it breaks. Real failures are rare and come from sequences of small things: a tilt reading that drifts, a command at the wrong moment, a slip. It is like finding the combination of a lock with many dials that must be turned in rhythm, 50 times a second.
What people did before
Push the robot with fixed or random forces; randomize the simulator during training so the controller copes on average; let engineers kick and pull it. These catch big shoves but don’t search for the subtle worst case. Earlier learned attacks targeted self-driving cars, or a simplified simulated quadruped without tests on a real robot.
What this paper does
It trains a second AI, an adversary, by trial and error to make the robot fall. Like a sparring partner that is not allowed to cheat, it may only do realistic things: bend the robot’s own tilt estimate by at most 3°, send velocity commands within ±0.5 m/s, and push the feet with up to 15 N in simulation. Its moves are kept smooth, like real sensor drift.
What they showed
Four standard push tests: 0 falls in 1000 trials. The adversary: a fall every time. The DARPA controller fell on flat ground with tilt errors of 3°, about what its tilt estimate already gets wrong on a plain floor, both in simulation and on the real ANYmal. Of 100 roboticists with joysticks, 3 managed a fall; the adversary was about 3 times faster than the quickest of them.
Why it’s a step forward
Robustness becomes something you search for rather than hope for. Training against the adversary produced a controller that new attackers could not topple even at 5°, on hardware too, while it walked just as well. Limits: only the robot’s body sensors were attacked (not its terrain map), and no method can promise to find every weakness.
- Controller (policy)
- the neural network that turns sensor readings into leg motions
- Reinforcement learning
- learning by trial and error, guided by a score
- Adversary
- a second AI trained to make the controller fail
- State estimator
- software that works out the robot’s tilt from its motion sensor and legs
- Domain randomization
- randomly varying pushes, friction and noise in simulation during training
- Fine-tuning
- continuing to train an existing controller on new situations