A four-legged robot with wheels for feet learned in simulation how to move and where to go, then drove itself for kilometres through Zurich and Seville.
Why this matters
City deliveries need robots that are quick on pavements but can also handle kerbs, stairs and doorways. Wheeled robots stop at the first staircase. Legged robots climb, but ANYmal averages about 2.2 km/h, half a walking pace, and runs for about an hour.
What makes it hard
With wheels on its legs, a robot can roll and step, but it must decide which to do at every moment, and no animal moves like that, so there is nothing to copy. At several metres per second it also has to react quickly. A plan that takes seconds is out of date before it arrives.
What people did before
Earlier controllers switched gaits with hand-written rules, fixed gait sequences or trajectory optimisation, which is slow and needs a good first guess. Navigation stacks drew cost maps and planned paths over several seconds. They assumed the robot would follow the path exactly, which led to pauses, zig-zags and near misses.
What this paper does
Two neural networks, stacked like a driver and a car. A locomotion policy takes a target speed and works out the gait itself. A navigator looks at the terrain, remembers where it has been and sets that target speed ten times a second. Both learned by trial and error in simulated worlds generated the way games generate levels.
What they showed
Kilometre-scale autonomous missions in Zurich and Seville, around pedestrians and up stairs. In Zurich it covered 8.3 km at 1.68 m/s on average, about three times a legged robot’s speed, using 53 % less mechanical effort per distance. Against a classic planner it was the only one with no collisions.
Why it’s a step forward
The navigator knows what its own legs can do, so gaits, speed and route choices fit together without hand-tuned glue. That makes wheeled-legged delivery robots practical. The limits: it sees only about 3 m ahead and understands shapes, not meanings such as “this is a lawn”.
- Reinforcement learning
- learning by trial and error in simulation, guided by a score
- Policy
- the trained neural network that decides what to do next
- Hierarchy
- a navigator policy giving speed targets to a locomotion policy
- Waypoint
- a point on the route the robot should head for next
- Cost of transport
- energy spent per unit of weight and distance; lower is more efficient
- Navigation graph
- a network of safe spots and connections, borrowed from video games