IEEE ICRA · 2023 · an animated walkthrough

Event-based Agile Object Catching with a Quadrupedal Robot

The paper in 98 seconds · narrated · sound on
Transcript

A ball flies at a four-legged robot, and the robot swings a net into its path. Its eye sees only what changes. How do you catch what you can barely see coming?

Legged robots can hike mountain trails, but a dog can catch a frisbee in mid-air. A ball thrown at 15 m/s crosses 4 m in about a quarter of a second.

Normal cameras must choose: few pictures per second, and the robot waits too long between looks; many, and a flood of data swamps its small computer. So the team gave ANYmal an event camera: every pixel speaks up the instant its brightness changes, and stays silent otherwise. With no frame clock to wait for, the robot knows the ball's path sooner.

First, it cancels its own motion, then keeps only the pixels that just changed. That's the ball. Next, it fits the ball's flight arc, ignoring bad sightings, and predicts where the ball will cross its body. A controller trained by trial and error in simulation then swings the net to that point. On the real robot, that means tilting, stepping or lunging to meet the ball.

Across real throws, it caught 83 % at up to 15 m/s, with all sensing and computing on board. It's the first quadruped to catch fast objects using only its own senses: a step toward robots that react like animals.

Robot footage: UZH Robotics and Perception Group (the paper's video; joint work with the Robotic Systems Lab, ETH Zurich). Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

Researchers in Zurich gave the four-legged robot ANYmal an eye that only reports what changes, so it can spot a ball thrown at up to 15 m/s and swing a net into its path, catching 83 % of throws with every calculation done on board.

  1. Why this matters

    Four-legged robots can already hike mountain trails using their cameras and laser scanners. Animals do more than that: a dog can snatch a frisbee out of the air. To handle a flying or falling object, a robot needs fast eyes and a fast body, and no quadruped had yet caught a fast object using only its own sensors.

  2. What makes it hard

    A ball at 15 m/s crosses 4 m in about a quarter of a second. A normal camera takes pictures on a fixed clock, and it needs at least two sightings to know where the ball is heading. Take few pictures and you wait too long between looks; take many and you drown the robot's small computer in data.

  3. What people did before

    Robot arms and drones have caught balls in rooms full of motion-capture cameras. A drone caught balls with its own camera, but only up to 6 m/s and with LED lights on the balls. Drones with event cameras dodged objects at up to 10 m/s, but dodging is easier than catching. Legged robots learned agile moves, but they saw the world through slow mapping steps.

  4. What this paper does

    It mounts an event camera on ANYmal: each pixel reports a brightness change the instant it happens, a bit like the motion-sensitive part of an animal's eye. The robot picks the ball out of this stream, fits its flight arc, and predicts where it will cross the robot. A controller trained by trial and error in simulation then throws the net to that point.

  5. What they showed

    With balls thrown by hand at 5–15 m/s from about 4 m, the robot caught 83 % of the throws that came within reach. It did this by tilting, stepping and lunging. The vision ran up to 100 times a second on a small onboard computer, and the controller moved from simulation to the real robot without extra tuning.

  6. Why it's a step forward

    It is the first quadruped to catch fast objects using only its own onboard sensing, and it shows that event cameras escape the camera's speed-versus-data trade-off on a legged robot. Honest limits: the net reaches about 0.6 m to either side, the single camera must know the ball's size, and steep, lobbed throws can be missed.

Words used below
Event camera
a sensor whose pixels each report only brightness changes, within microseconds
Latency
the delay between something happening and the robot knowing about it
Bandwidth
how much data per second the sensor sends and the computer must process
Parabola
the curved arc any thrown object follows under gravity
Impact point
where the ball's arc will cross the robot: the spot the net must reach
Reinforcement learning
training a controller by trial, error and reward, here in simulation
1 / 7
ball, its arc and impact point event camera and vision robot controller and net
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper (numbered as in the arXiv PDF); the text is a plain-language walkthrough.

Scene 1

Why catching matters

In short: legged robots can handle rough ground, but not yet things that come flying at them. Catching a fast ball is a sharp test of both quick seeing and quick moving.

The paper starts from animals. Humans and animals combine fast, precise perception with robust control; the authors' example is a dog catching a frisbee in mid-air. Quadrupedal robots have recently reached animal-like performance on hard terrain: the same group's ANYmal hiked in the Alps using its sensors, and ANYmal learned to get back up after being flipped over. Yet, the authors write, current quadrupeds are still far from the agility of cats and dogs, and catching fast objects with onboard sensing had not been done.

The test in this paper: balls are thrown by hand at up to 15 m/s (about 54 km/h) from about 4 m. By our simple arithmetic, 4 m at 15 m/s leaves the robot roughly 0.27 s. In that time it must see the ball, work out where it will arrive, and move a net mounted on its back there. The net is a landing net 0.3 m across, whose centre sits 0.5 m above the robot's body.

Paper Fig. 1. What the event camera sees (left, red and blue dots are brightness changes), the camera itself on the robot (middle), and ANYmal catching a ball (right).
Scene 2

Frames vs. events

In short: normal cameras and laser scanners force a choice between reacting late and producing too much data. Event cameras avoid that choice.

Robots usually see with standard (RGB) cameras or LiDARs. The paper's key point is that these sensors suffer from a bandwidth–latency trade-off. To react quickly you need a high frame rate, but every frame is a full picture, so the data rate climbs. Lower the frame rate and the data is manageable, but the robot waits longer for its next look, which hurts its reaction time and success rate. On a robot with limited onboard computing, you cannot have both.

An event camera is a “neuromorphic” sensor, meaning it copies how biological eyes work. It has no frames at all. Each pixel independently sends a tiny message, called an event, whenever its brightness changes. The event holds the pixel's position, a timestamp with microsecond precision, and the sign of the change (brighter or darker). A still scene sends almost nothing, while a moving ball lights up only the pixels along its edges. The robot here uses an iniVation DVXplorer with VGA resolution (640 × 480 pixels) and a 90° lens.

Animated. A normal camera samples this pixel on every clock tick, even when nothing happens. The event pixel stays silent until the ball arrives, then fires brighter (red) and darker (blue) events right as the change happens. Illustrative.

Event cameras had already been used for fast dodging on drones. The open question the paper takes on is whether they can serve a very different robot, a heavy legged one, for the harder task of catching.

Scene 3

The latency race

In short: the paper works out the worst-case delay before a robot knows a ball's path. With events, the delay depends only on how fast the detector runs, not on a camera clock.

To predict a flight arc from positions alone you need at least two 3-D sightings; the paper calls these two-shot methods. The analysis extends earlier work by Falanga et al. and asks: in the worst case, how long until the second sighting is done?

Standard camera with frame interval ΔtFPS and detector time Δtc: the ball may appear just after a frame was taken, so the robot waits for the next frame, detects, waits for the following frame, and detects again.

Δtl = 2·ΔtFPS + Δtc  if Δtc < ΔtFPS, otherwise 3·Δtc  // worst-case latency, normal camera
Δtl = 3·Δtc  // event camera: no frame clock, only the detector's own cycle

For events, ΔtFPS is effectively zero, so the only wait is for a detection cycle that was already running. The event-based method therefore has lower latency whenever one detection is faster than one frame interval (Δtc < ΔtFPS). Events also carry motion information, so a future one-shot method that reads position and speed at once could get down to 2·Δtc. The paper's own method is two-shot and leaves one-shot methods for future work.

Animated. The paper's two formulas, plotted for an example 30 fps camera. As long as detection takes less than one frame (33 ms here), the event line stays below the camera line. The 30 fps value is our example, not a measurement.
Paper Fig. 3. Timelines of the worst case for a standard camera (waiting for frames) and an event camera (waiting only for computation), and the resulting latency curves. The red arrow marks the latency gain.
Scene 4

Spotting the ball

In short: the robot's own motion makes the whole room fire events. The method cancels that motion, then finds the pixels whose events are all recent: that is the ball.

The whole pipeline has two halves: a visual front end that turns events into a predicted impact point, and a control back end that moves the robot. The front end builds on the moving-object detector that the Robotics and Perception Group used for drones dodging objects (Falanga et al., Science Robotics 2020).

1 · Cancel the robot's own rotation

The detector keeps a rolling buffer of recent events. Because the robot moves, background edges also produce events. The IMU (the robot's motion sensor, which measures how fast it turns) gives the average rotation rate over the buffer, and each event is shifted back to where it would have landed without that rotation. Events caused only by the robot's turning then line up on the still edges.

2 · The mean timestamp image

For every pixel, the method averages the timestamps of the events that landed there. Background pixels, which the rotation explains, collect events across the whole buffer, so their average is unremarkable. Pixels on an independently moving object stand out with higher values. After normalising, a threshold keeps only those pixels. The threshold rises with the rotation rate, since faster turning raises the background values too.

keep pixel x if ρ(x) > θ0 + θ1·‖ω̄‖  // ρ: normalised mean timestamp; ω̄: mean rotation rate from the IMU

3 · Group pixels into objects

The kept pixels are grouped with DBSCAN, a clustering method, using three kinds of similarity: how close pixels are, how similarly they move (optical flow computed between consecutive timestamp images), and how similar their timestamp values are. Each cluster gets a bounding box.

4 · Distance from size

With one camera, depth is not measured directly. The method assumes the object's real size is known (the balls are 10 cm across) and uses the pinhole-camera rule Z = f · Wreal / Wimage. The box centre gives the direction, the box width gives the distance, and the robot's odometry turns the result into a 3-D point in a fixed world frame.

Animated. A 10 cm ball at 1 m fills about 32 pixels; at 4 m, about 8. A pixel or two of error in the box width therefore moves the distance estimate a lot, which is why these 3-D points are noisiest along the camera's line of sight. The focal length is our estimate from the camera's specs.
Paper Fig. 6a. Raw events seen from ANYmal during an experiment. Red and blue mark brightening and darkening; the whole cluttered lab shows up, not only the ball.
Paper Fig. 4. The whole method on one page: events → mean timestamp image → clusters → 3-D points and a fitted arc (green) → impact point on the plane at the robot → the policy network → joint commands.
Scene 5

Fitting the flight arc

In short: a thrown ball follows a predictable arc. The method fits that arc to noisy sightings while ignoring wrong ones, then computes where the arc crosses the robot.

A ball in flight follows a parabola set by its starting position and velocity, with gravity known:

p(t) = p0 + v0·t + ½·g·t²  // g = (0, 0, −9.81 m/s²) in the world frame

Because gravity is known, two sightings are enough to fix p0 and v0. Some sightings are simply wrong, so the method uses RANSAC: it repeatedly picks a random pair, builds the arc through it, and counts how many other sightings lie close to that arc. Closeness is measured with a stretched distance: errors along the camera's depth direction count less (variance 5 versus 1 for the other directions), because depth-from-size is the noisiest part. The pair with the most agreeing points wins, and a final least-squares fit on all of those points gives the arc. This fit has a closed-form solution, so it is fast.

Animated. An arc through a pair that includes a bad sighting convinces few points. An arc through two good sightings wins the vote. Illustrative points.

Where will it cross?

The controller needs one target: the impact point. The method intersects the arc with an upright plane through the robot's body, spanned by gravity and the robot's sideways axis. It solves for the crossing time timp and evaluates the arc there. A median filter smooths the stream of impact-point estimates until the robot starts its catching move. With this setup, the vision system runs at up to 100 Hz on an NVIDIA Jetson Orin.

timp = −(n·p0) / (n·v0),   pimp = p(timp)  // n: normal of the plane at the robot, in the body frame
Real run. The robot's own view of one throw, replayed as a 3-D visualization: red dots are ball detections, the green curve is the fitted flight arc, and the blue sphere near the robot marks where the arc is predicted to cross it. Footage: UZH Robotics and Perception Group, from the paper's video (joint work with the Robotic Systems Lab, ETH Zurich).
Paper Fig. 6b. In a real experiment: 3-D detections, the fitted arc (dark green) and the impact point (blue circle) the robot will aim its net at.
Scene 6

The learned catcher

In short: instead of hand-planning every move, the team let a neural network learn to put the net on a target point, by practising in simulation.

Planning whole-body motions for a legged robot with physics models is computationally heavy, which makes fast reactions hard. Recent work showed that deep reinforcement learning can produce agile, real-time controllers. Here the task is simple to state: reach the impact point with the net as fast as possible. It is trained with PPO (a standard reinforcement-learning algorithm) in Isaac Gym, a GPU simulator that runs many robots in parallel. The simulated ANYmal C uses an “actuator network” that models the real motors.

What the controller sees and does

  • Input (48 numbers): the impact point in the robot's body frame (3), plus the robot's body sense (45): gravity direction, body linear and angular velocity, the 12 joint angles and their speeds, and its previous action.
  • Output: for each of the 12 joints, how far to move from its standing angle.

How it practised

In simulation, a target net position is sampled around the current net position: within ±30 cm forward/back, ±80 cm sideways and ±15 cm up/down. Every second it jumps to a new place with probability 0.2. To make the policy robust, the training adds sensor noise, slightly rough ground, random pushes, and random friction and mass. An episode lasts up to 20 s and ends if the body or a thigh hits the ground.

r = 4·exp(−10·‖pimp − pnet‖²) − small penalties  // penalties: vertical bouncing, body wobble, joint speed, acceleration, torque, jerky actions, body tilt
Animated. The main reward is the paper's formula: 4 when the net is on target, falling off quickly with distance (about 1.6 at 30 cm). The robot's motion is illustrative.
Paper Fig. 5. The training world: many simulated ANYmals, each trying to put its net on a yellow target sphere while being pushed around on slightly rough ground.

The trained policy was deployed straight onto the real ANYmal, and the authors found it needed no additional fine-tuning. It runs on the robot's onboard PC (Intel Core i7-8850H), while the vision runs on the Jetson Orin. A laptop was connected only to show debugging views; all computation for catching ran on board.

Real run. The same learned controller on the real ANYmal, three catches side by side: it tilts its body, takes a step, or lunges, depending on where the ball will arrive. Footage: UZH Robotics and Perception Group, from the paper's video (joint work with the Robotic Systems Lab, ETH Zurich).
Paper Fig. 7. Two successful catches, frame by frame. The robot has to tilt sharply and walk sideways. The fastest ball speeds in the two sequences were 11.5 m/s and 10 m/s.
Scene 7

Results

In short: on the real robot, 83 % of throws up to 15 m/s were caught. The vision is very reliable near the net and weaker for fast throws far to the side.

Setup. People threw balls by hand at the robot in a cluttered lab. Two different rubber balls, each 10 cm across, tested whether the method copes with different surface textures. Throws averaged 5–15 m/s, landed up to about 0.6 m from the net's centre, and arrived at roughly 90 ± 15° to the net. 22 throws came within the 0.6 m working range. Beyond that range ANYmal would have to walk sideways instead of lunging, which is slower. True ball speeds and impact points came from a 60 fps external camera.

Real run. A hand throw in the cluttered lab, slowed to a quarter of real speed. The green grid in the authors' video marks the area where the robot can catch. Footage: UZH Robotics and Perception Group, from the paper's video (joint work with the Robotic Systems Lab, ETH Zurich).
Throw speed< 8 m/s< 10 m/s< 12 m/s< 15 m/s
Catch success (Table I)81 %88 %92 %83 %

What this means: across every speed group, more than four out of five balls ended up in the net. The best result, 92 %, was just below 12 m/s (about 43 km/h).

Distance from net centrethrows < 9 m/sthrows < 15 m/s
< 0.4 m93.33 %88.88 %
< 0.6 m73.68 %50.25 %

What this means: Table II scores the vision alone. It shows how often the predicted impact point landed within one net radius of the truth. Close to the net it is right about nine times in ten. For fast throws far to the side it is right only half the time.

Why performance drops at the edges. The vision degrades above about 9 m/s and beyond 0.4 m. Faster balls give fewer sightings to fit the arc, and balls thrown far out leave the camera's field of view sooner. Catching has a sweet spot. Too fast, and vision fails more often and the robot cannot move in time. Too slow, and throws vary more in range and come in at steeper angles, which shrinks the part of the net the ball can enter. The policy is never told the arrival angle.

Animated. Seen from a ball arriving at an angle, the round net looks like a narrower ellipse, so there is less room to get in. This is simple geometry that illustrates the paper's explanation.
Paper Fig. 6c. The catching manoeuvre on the real robot.
Wrap-up

Limits and next steps

In short: this is a first working demonstration, and the authors list clear ways to extend its reach and reliability.

  • Reach. Catches work for balls arriving within about 0.4 m of the net's centre; the vision stays reliable out to 0.6 m. A camera with a wider field of view could extend the robot's range to the side.
  • One camera needs a known size. Depth comes from the ball's known width. A stereo (two-camera) setup would remove that assumption and reduce depth noise, so less filtering would be needed.
  • A smarter controller. The policy simply reaches the target as fast as possible. Training it with the ball's motion, the sensor's behaviour and the detection pipeline in the loop could improve it further. Giving it the ball's velocity at impact could fix the misses caused by steep arrivals.
  • One-shot perception. Events carry motion information, so a method that measures position and velocity from a single detection could cut the worst-case delay from 3·Δtc to 2·Δtc.

The conclusion ties it back to the story. Common robot sensors are caught in a trade-off between delay and data volume. An event camera breaks that cycle, and combined with a learned controller it gave a quadruped the ability to catch objects thrown at up to 15 m/s.