IEEE ICRA · 2021 · an animated walkthrough

Circus ANYmal: A Quadruped Learning Dexterous Manipulation with Its Limbs

The paper in 100 seconds · narrated · sound on
Transcript

This walking robot is lying on its back, with a yoga ball on its feet. It turns the ball with its legs, shrugs off a poke, and a push on its leg. It learned all this in simulation. How?

Legged robots now walk almost anywhere. But mostly they only look. To handle things, they carry an arm, and there’s room for about one. Animals do it differently: a cat on its back bats a toy with all four paws. Could a robot’s legs work like fingers?

That’s hard. The feet must keep grabbing and releasing a soft ball — and the real ball squashes and slips, unlike a simulator’s rigid sphere. Lying on its back, the robot is asked to turn the ball about any axis: roll, pitch or yaw.

Instead of rewarding speed, the goal is a target orientation that steps ahead three times a second. A neural network — the policy — learns by trial and error in simulation, called reinforcement learning, while the ball’s size, weight and grip keep changing, and it gets shoved.

Then the same network ran on the real robot, unchanged, turning the ball at 15 °/s. With a 3 kg yoga ball, it kept going for over 2 minutes, and recovered when people poked the ball. Most of the time, it used under 5 of its motors’ 80 newton metres.

Flip the picture, and walking over a hill becomes turning a ball. This was the first real quadruped to handle an object this dexterously with its legs.

Real-robot footage: Robotic Systems Lab, ETH Zurich (the paper’s experiment video). Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

A four-legged walking robot learned to lie on its back and turn a yoga ball with its feet, like a cat playing with a toy — the first real quadruped to handle an object this dexterously with its limbs.

  1. Why this matters

    Legged robots such as ANYmal can run, get up after a fall and cross rough ground, so they already go on inspection missions. But mostly they look. They rarely touch or move anything.

  2. What makes it hard

    The usual fix, bolting on an arm, is limited: these robots can carry about one. Using the legs instead means juggling four feet on a soft, bouncy, slippery ball — like balancing a beach ball on your fingertips while it squashes and slides.

  3. What people did before

    Robots have lifted a box with two legs while standing on the other two, and a simulated quadruped balanced on a ball. Robot hands trained by trial and error learned to turn blocks and even solve a Rubik’s Cube. No real quadruped had handled an object dexterously with its legs.

  4. What this paper does

    Treat the legs as the fingers of a hand. A neural network learns by trial and error in a physics simulator, scored on keeping the ball close to a target orientation that steps ahead three times a second. The simulator keeps changing, so the real world is just one more variation.

  5. What they showed

    The same network ran on the real 40 kg robot unchanged and turned a 3 kg, 0.8 m yoga ball about all three axes at up to 15 °/s — a full turn every 24 seconds. It ran for over 2 minutes and recovered when people poked the ball or a leg.

  6. Why it’s a step forward

    To the authors’ knowledge, it is the first dexterous, dynamic object handling on a real quadruped — with no touch sensors. It also tests an old idea, that walking and in-hand manipulation are two sides of one problem. Limits: ball pose came from motion capture, and above 20 °/s the feet slip more.

Words used below
Reinforcement learning
learning by trial and error from a score (the reward).
Policy
the neural network that turns sensor readings into joint commands.
Domain randomization
randomly varying the simulator so reality is just another variant.
Zero-shot transfer
running the simulation-trained policy on the real robot with no retuning.
Roll · pitch · yaw
the three ways to turn an object: sideways, forward, spin.
1 / 8
Target / command Simulation Learned policy
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough.

Scene 1

Robots that walk, but rarely handle

In short: four-legged robots have become excellent walkers, but to do useful work they also need to handle things — and the paper asks whether their legs could do that job, the way an animal’s do.

By 2020, quadrupeds could run fast, recover from falls and walk over stairs and rough ground, and ANYmal had been used on real inspection and exploration missions. Most of those missions are about moving and looking. Actively touching and moving things was still largely missing.

Several groups gave their quadrupeds an arm and gripper, which enables pick-and-place, opening doors or carrying things together. That works, but it treats handling as a separate machine on top of walking, and the robots’ payload usually allows only one arm.

Animals do something different. A cat lying on its back bats and turns a toy with all four paws; a baby grabs toys with hands and feet. The paper takes that as its inspiration: use the limbs the robot already has.

Paper Fig. 1. The inspiration — a cat (a) and a baby (b) handling things with their limbs — and the result: ANYmal on its back holding a yoga ball on its feet (c).

Earlier work using legs as hands

  • Two legs lift a box while the other two keep the robot standing (quadrupeds); on six-legged robots, four legs stand and two manipulate.
  • Extra support legs added to ANYmal to make combined walking-and-handling more robust.
  • An inchworm-like motion with all four legs to move two boards.
  • A quadruped balancing on a ball and moving it at the same time — in simulation, with a model-based controller.

What was missing: a real quadruped using all four limbs for dexterous, dynamic handling — continuously turning an object, not just lifting it.

Scene 2

Why a soft ball is hard

In short: turning a ball with four feet means constantly making and breaking contacts that are too messy to model exactly — and the real ball differs from the simulated one in several ways.

Legged walking and in-hand manipulation have long been called a dual problem: legs are like the fingers of a hand, just pushing on the ground instead of on an object. In-hand manipulation is famously hard because a many-jointed hand has to handle contact that is difficult to model accurately and expensive to compute in real time.

Deep reinforcement learning had recently made big progress there — robot hands learned to operate a valve, rotate a block and even solve a Rubik’s Cube — and, separately, on quadruped walking, including ANYmal. This paper brings the two lines together.

The paper names the gaps between its simulator and the real experiment:

  • Softness and deformation. The simulator (Raisim) treats bodies as rigid, but a yoga ball dents under a foot. Contact forces in the pressing direction differ a lot between simulation and reality.
  • Slip between the feet and the rubber surface.
  • Shape and centre of mass. After inflation the ball is not a perfect sphere, and its centre of mass is off.
  • Joint position tracking errors on the robot.

On top of that, the legs are strong for the job. Each of the 12 series elastic actuators (motors with a spring in series) gives up to 80 Nm, while the ball weighs 3 kg. Robot-hand fingers are the opposite — small motors, limited torque — which is why hand papers could reward only the orientation. Here the reward also has to rein in effort.

Scene 3

Legs as fingers

In short: the robot lies on its back, the ball sits on its feet, and a user asks for a turning speed about one axis; only the feet are allowed to touch the ball.

The robot is ANYmal-C: about 40 kg, with three motors per leg — hip sideways (HAA), hip forward-back (HFE) and knee (KFE). During the task it lies on its back, so the usual body position and velocity are left out of what it senses. The base must not touch the ball.

The user gives a target angular velocity for the ball. To show that the robot can turn the ball in any direction in 3-D, the policy is trained to turn it about each of the three axes: roll, pitch and yaw.

To keep the problem manageable, two things are given rather than sensed by the robot: the ball’s position and orientation come from an external motion-capture system (VICON), and the ball’s mass and radius are known. The robot feels its own joints only — no contact sensors, no force sensors.

Paper Fig. 2. The task in the Raisim simulator. Each leg has three joints (arrows), driven by series elastic actuators.

The 3-D animation places the ball centre about 1 m above the floor, matching the measured ball height in Fig. 5b; limb proportions and foot placement are approximate.

Scene 4

A target that steps forward

In short: instead of rewarding the ball’s speed, the robot is rewarded for keeping the ball’s orientation close to a target that moves ahead in small steps — plus a few penalties that keep the motion gentle.

Walking policies are usually rewarded for matching a commanded velocity. Here the command is converted into two parts that are updated periodically: (1) a target orientation — where the ball should be facing next — and (2) the time left to reach it. The authors give two reasons: holding a constant speed on a round, changing “terrain” is hard (the same was seen for rough-terrain walking), and for most manipulation tasks it is accurate orientation that matters, not constant speed. In training, this design converged better than a direct velocity reward.

Animated. The same 15 °/s command becomes a staircase of targets. Training moved from 1 Hz to 2 Hz to the final 3 Hz (every 0.33 s).

The reward

The main term scores the angle δq between the ball’s current and target orientation with a bell-shaped function that peaks at zero error:

rq = kq / (eδq + 2 + e−δq)  // δq = 2·acos(q0), from the quaternion difference to the target
Animated. As the error shrinks, the dot climbs to the top of the bell. The curve’s scale kq is set to 1 here.

Four penalties keep the behaviour safe and realistic:

Penalty onWhy
Robot body velocityAggressive leg motion shakes the base, which risks the hardware. Dominant early in training, fades to zero by the middle.
Joint torque (squared)Keeps effort low and stops the legs from squeezing the ball, which can deform it and fail the task.
Foot slip (sideways speed at contact)Only moderate: forbidding slip entirely is unrealistic and would widen the sim-to-real gap — a limitation of an earlier model-based ball-balancing method.
Foot speed into the ballThe simulator can’t deform the ball; punishing fast pushes into it avoids relying on contact forces that would differ on the soft yoga ball.

What this means: the reward is intentionally small — one “do the task” term and four “do it gently” terms — and two usual walking terms (joint-speed and foot-clearance rewards) were dropped.

Cutting a practice run short

Ending a run early when something goes wrong steers learning away from strange solutions. A run is stopped when:

Animated. The four stop rules. The allowed box is ±1.5 ball radii horizontally and ±1 radius vertically; the no-contact limit stops the robot from tossing the ball into the air.
Scene 5

Practice in a simulator that keeps changing

In short: the policy practises in simulation where the ball, the robot and the sensors are randomly altered and the ball is shoved, and the task gets harder as training goes on.

Everything is learned in the Raisim simulator with model-free reinforcement learning. To make the motors behave realistically, simulation uses an actuator network — a small learned model of ANYmal’s series elastic actuators from earlier RSL work.

Domain randomization

What is variedHow much (paper, §II-E)
Leg geometry (shank positions and lengths)noise 0.03 m
Joint readings0.05 rad (angle), 0.3 rad/s (speed)
Ball mass / radius±5 % / ±10 %
Ball friction / restitution (bounciness)uniform 0.5–1.1 / 0.9–1.0
Ball position / orientation readings0.04 m / 0.03 rad per axis
Starting pose of robot and ballrandom

What this means: the policy never sees the same ball or the same robot twice, so the real, imperfect yoga ball is just one more variation it has learned to handle. (The paper writes the noise terms as N(0.03, 0) etc.; we quote the numbers as given.)

Paper Fig. 3. Randomization of the ball and robot models and sensor readings, and an external push on the ball during training.
Simulation. Nine practice runs side by side, each with its own noise, model changes and shoves; the small ball beside each robot shows the target orientation. Footage: Robotic Systems Lab, ETH Zurich (authors’ video).

Random shoves

Pushing a robot around in simulation had already helped walking policies transfer to reality. Here the ball receives a 50 N force from a random direction, lasting 0.4 s, at random moments with 20 % probability.

Animated. A shove knocks the ball sideways; the feet catch it.

Curriculum

Training starts easy and gets harder. The target speed rises from 0 to 15 °/s, randomization is scaled by a factor that grows to 1, and the target update period shrinks from 1 s to 0.5 s to 0.33 s (1, 2 and 3 Hz). With a longer period the motion is smoother, with fewer contact switches, but fails more easily once speed and noise go up. A shorter period means more contact switching but allows higher speed and more robustness. Joint speed and acceleration limits rule out going much faster, so — like choosing a gait for walking — the authors settled on 3 Hz.

They also saw the policy become more cautious as disturbances and randomization grew: for example, the feet were lifted less, so they could react to surprises faster.

The paper gives the start and end values of the curriculum but not its exact schedule; the ramps in the animation are illustrative.

Scene 6

The policy

In short: a small neural network reads the joints, the ball and the target, and outputs 12 joint-angle targets 100 times a second.

Each moment, the policy gets 43 numbers: 12 joint angles, 12 joint speeds, the ball’s position (3), the quaternion difference between the ball’s orientation and the target (4), and its own previous 12-number command. It gets these for now, 10 ms ago and 20 ms ago — the same history length the actuator network uses — which helps it deal with delays and things it can’t observe directly. With the time left to reach the target, that makes 130 inputs.

xt = (θt, θ̇t, pt, qt, at)    st = (xt, xt−1, xt−2, tremain) ∈ ℝ130

The network is a multi-layer perceptron with two hidden layers of 256 and 128 units and tanh activations. Its 12 outputs are joint position targets that each joint’s PD controller turns into motor torque. A second network of the same shape estimates how good a situation is (the “value”) for training. Training uses PPO (proximal policy optimization) with discount factor 0.998, clipping 0.2 and learning rate 0.001.

Paper Fig. 4. The policy network: 130 inputs, hidden layers of 256 and 128, and 12 joint position commands out.
SettingValue
Inputs130
Hidden layers256, 128 (tanh)
Outputs12 joint targets
Rate on the robot100 Hz
PPO: γ / clip / learning rate0.998 / 0.2 / 0.001

What this means: the “brain” is tiny by today’s standards; the difficulty lies in what it is trained on, not in its size.

Scene 7

On the real robot

In short: without any adjustment, the simulation-trained policy turned a real yoga ball about all three axes, kept going for minutes and recovered from pokes, using only a small fraction of the motors’ strength.

The setup: ANYmal-C (about 40 kg, 12 actuators with up to 80 Nm), a commercial yoga ball of 3 kg and 0.8 m diameter, and a VICON motion-capture system tracking the ball. The policy ran at 100 Hz, sending joint targets to the robot’s joint controllers. Roll, pitch and yaw were tested separately — turning freely in 3-D would have let the feet hit the tracking markers on the ball — each at 10 and 15 °/s.

Result on hardwareValue
Top turning speed15 °/s
Axes demonstratedroll, pitch, yaw (separately)
Continuous run> 2 min
Torque used, most of the time< 5 Nm of 80 Nm
Peak and mean joint torque and speed≈ ¼ of walking

What this means: 15 °/s is one full turn of the ball every 24 seconds — slow for a human, but it is a heavy robot doing it by feel, without touch sensors, and never trained on the real ball.

Paper Fig. 5 (a, b). Yaw at 15 °/s over 50 s. Left: the four quaternion numbers describing the ball’s orientation, target (dashed) vs. measured — nearly on top of each other. Right: the ball’s position stays put, while its measured turning speed jitters around the target.
Paper Fig. 5 (c, d). Pitch at 10 °/s with two pokes (marked): small bumps in the tracking, then back on target. Right: joint torques and speeds stay small; the paper notes torque is below 5 Nm most of the time.
Animated. How little of the motors’ strength the task needs.

Pokes

People poked the ball and the robot’s legs with a stick from different directions. There is no high-level planner and the grip is non-prehensile (the feet push, they don’t grasp), which is inherently less stable — yet the robot quickly recovered and carried on.

Paper Fig. 6. Rolling at 15 °/s: the ball is poked and visibly dents (a), is knocked off (b), and the policy recovers (c).
Real run. Rolling at 15 °/s while the ball is poked with a stick from the side; the legs keep it on top and carry on. Footage: Robotic Systems Lab, ETH Zurich (authors’ video).
Paper Fig. 7. Pitching at 10 °/s: a leg is pushed so its foot moves (a, b), the ball’s orientation swings (c), and the legs bring it back (d–f).
Real run. Pitching at 10 °/s while a stick pushes one of the legs; the ball keeps turning. Footage: Robotic Systems Lab, ETH Zurich (authors’ video).

The paper is also candid about the fluctuations: even when the orientation is tracked well, the turning speed fluctuates around the target (Fig. 5b).

Scene 8

Walking ≈ handling

In short: flip the picture and walking on a round hill becomes turning a ball — the paper compares what the two problems share and where they differ when both are learned.

From learned walking, the paper borrows a short sensing history, random pushes during training and a curriculum. But several standard walking ingredients did not fit:

IngredientTypical in learned walkingIn this ball task
Speed rewardreward velocity directlytarget orientation stepping ahead; converged better
Foot clearancereward lifting feet to a chosen heightremoved — no intuitive height on a curved object
Gait patternoften a preset contact sequencenone given; no animal gait to copy for in-limb handling
Model-based controllearned gait + whole-body controller worksunsuitable: deformation and slip defeat accurate models
Motor strength(robot hands) small finger motors — orientation reward alone sufficesstrong legs → effort must be penalised

What this means: the “duality” is real enough that walking tools help, but handling needs its own reward design.

The authors also watched people turn a soccer ball with their fingers and noticed similar undulating motions — while noting that a direct comparison isn’t possible.

Wrap-up

Limits and what’s next

In short: a first step — slow, one axis at a time, with the ball tracked from outside — but a real one.

  • Speed. Above 20 °/s the feet slip more (legs with few joints and a small contact area), making failure likelier; faster, harder leg motion also pushes the robot’s base hard against the floor, which risks damage.
  • External sensing. Ball position and orientation came from a motion-capture system, and the ball’s mass and radius were known.
  • One axis at a time. Roll, pitch and yaw were shown separately on hardware, to keep the feet off the tracking markers.
  • Non-prehensile. The feet push rather than grasp, and there is no high-level planner.

The authors’ closing hope is that comparing reward design for walking and handling will bring the two fields closer together.