IEEE ICRA · 2025 · an animated walkthrough

Learning Quiet Walking for a Small Home Robot

The paper in 105 seconds · narrated · sound on
Transcript

Robot pets are moving into our homes, but their footsteps can be loud: 35.9 dB for a standard learned walk. A new walk, learned in simulation, measures just 22.7 dB, quieter than Sony’s own quiet mode. So how do you teach a robot dog to tiptoe?

Sony’s aibo is a robot dog that lives with people. On a hard floor every step clacks, and owners say the walking sound is too loud. Robots now learn to walk by trial and error in a simulator. But a simulator can’t hear. So the team used something it can measure: how fast the paw hits the floor. A slower touchdown carries less energy, and makes a quieter step.

Trick one: the network also sets how stiff each joint is. Soft as the paw lands, firm when the leg carries weight. Trick two: aibo has no force sensors, only a simple switch in each paw. It tells the robot the instant a paw lands. Trick three: punish noise from the start and the robot just stops walking. So it first learns to walk, then the noise penalty is raised five-fold.

On the real robot: 22.7 dB, 13 dB below a standard learned walk, and below both of Sony’s hand-tuned controllers. The price: the quietest walk climbs only a 3° slope; the loud one, 7°. Bumpier training floors win that back, for a little more noise. Here is that bumpier-floor version on a real 7° ramp. It walks right up. The lesson: if a simulator can’t model what you care about, minimize something it can compute that goes with it.

Learning Quiet Walking for a Small Home Robot.

Real footage (opening and ramp clip): the authors’ result videos from the QuietWalk project page (Sony Group Corporation and the Robotic Systems Lab, ETH Zurich). Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

Robot dogs like Sony’s aibo clatter across the floor. The researchers taught aibo, inside a simulator, to set each paw down softly, and on the real robot it walked more quietly than Sony’s own hand-tuned “quiet” mode.

  1. Why this matters

    Home robots like Sony’s aibo robot dog live with people as companions. But its hard paws clack on the floor, and the paper reports that one of the main concerns of aibo owners is that its walking sound is too loud.

  2. What makes it hard

    Robots now learn to walk by trial and error in a physics simulator, and a simulator can’t hear. aibo is an affordable robot with no force sensors in its legs. And if you simply punish noise, the robot finds a shortcut: it stops walking.

  3. What people did before

    Learned gaits made legged robots very robust on rough ground, and some made them energy-efficient, but none aimed at being quiet. Energy-efficient gaits are not automatically quiet. Sound design had been studied for robot arms, two-legged and rolling robots, but not for four-legged footsteps.

  4. What this paper does

    It trains the robot to land its paws slowly, because touchdown speed is something the simulator can measure and it goes together with footstep sound. Think of putting a full cup on the table without a clink: slow down just before contact and relax your wrist. The robot learns to soften its joints, uses its paw switches to feel the floor, and learns to walk before it learns to walk quietly.

  5. What they showed

    On the real aibo, recorded by its own head microphone, the new walk measured 22.7 dB against 35.9 dB for a standard learned walk: 13 dB quieter, about a twentieth of the sound power. It was also quieter than both of Sony’s commercial controllers, including their “quiet” one, at every speed tested.

  6. Why it’s a step forward

    It is the first work to show which ingredients a learned controller needs to walk quietly on a real robot. The trick generalizes: if you can’t simulate what you care about, minimize something the simulator can compute that tracks it. The honest cost: the quietest walk is less sure-footed on slopes.

Words used below
reinforcement learning
learning by trial and error, rewarded for good behaviour; here in a simulator
foot contact velocity
how fast a paw is moving at the instant it touches the floor
PD gains
how stiff (P) and how damped (D) a joint motor acts
curriculum
training in stages: first walking, then walking quietly
domain randomization
varying floor, friction and weight during training so the walk copes with reality
decibel (dB)
a log scale of loudness: 10 dB less means one tenth of the sound power
1 / 7
stiff / loud soft / damped the quiet policy
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough.

Scene 1

A robot pet that is too loud

In short: robot pets now live in homes, and for them a quiet walk matters as much as a sturdy one. Nobody had trained a legged robot to walk quietly.

Home robots are spreading as actuators, sensors and AI improve. Sony’s aibo, a small robot dog, develops its own personality through life with its owner, and studies cited in the paper show people form real emotional bonds with it. But a home is not a lab. The paper reports that one of the main concerns of aibo users is that the walking sound is too loud, mostly the clack of each paw hitting the floor.

Over the last few years, reinforcement learning (a robot learning by trial and error in simulation, rewarded for good behaviour) has made four-legged robots remarkably robust on stairs, rubble and snow. That research optimized for robustness and efficiency. Nobody had optimized for quiet. Even an energy-efficient learned gait (Yang et al.) did not turn out quieter. Research on robot sound had looked at arms, two-legged walkers, rolling robots and servomotors, but not at four-legged footsteps.

Paper Fig. 1. The idea in one picture: in the simulator (left) the policy minimizes how fast the paw hits the ground; the same policy runs unchanged on the real aibo (right), where that shows up as a quieter footstep.
Scene 2

A stand-in for sound

In short: simulators can’t produce sound, so the robot is trained to reduce the paw’s speed at touchdown, which goes together with how loud the step is.

Training happens in Isaac Gym, a GPU physics simulator in which many virtual aibos practice in parallel, with physics at 400 Hz and the controller at 100 Hz. The simulator knows every position and speed, but it has no acoustics. The authors therefore pick a proxy, a quantity the simulator can compute that tracks the real thing.

Biomechanics research shows that footstep sound goes up with foot contact velocity, the paw’s speed at the moment it touches the floor. At impact the paw’s kinetic energy, ½mv², turns into other forms of energy, and some of it becomes sound. Two more quantities support the proxy: joint acceleration and body angular acceleration. The authors observed that sudden jerks of the joints and body tend to follow loud steps, because a noisy walk bumps into the ground.

rcontact = −w · ‖vfoot,xyz‖²  // penalty on paw speed while the paw is in contact; weight w = 5 → 25 (Table II)
Animated. The two paws land at the paper’s average touchdown speeds in simulation. Kinetic energy grows with speed squared, so the quiet gait lands with about 1/11 of the energy. The speed profile before touchdown is illustrative.
In simulation (Table III)Standard RL gaitProposed
Foot contact velocity0.417 m/s0.123 m/s
Joint acceleration114.3 rad/s²76.7 rad/s²
Body angular acceleration57.2 rad/s²23.7 rad/s²

What this means: the policy that is quieter on the real robot also scores lower on all three simulated quantities (averaged over 10 s), which supports using them as stand-ins for sound. The paw lands about 3.4 times slower.

Scene 3

Stiff or soft

In short: the network decides not only where each joint should go but also how stiff it should be. It can relax a leg just as the paw lands and firm it up to carry weight.

aibo has 12 leg joints, three per leg: shoulder roll, shoulder pitch and ankle pitch. Each is driven by a PD controller, which works like a spring and a shock absorber. The P gain is the spring pulling the joint towards its target angle; the D gain is the damper resisting fast motion. Most learned controllers output only target angles and keep the gains fixed.

Here the policy outputs 24 numbers every 10 ms: 12 target angles and 12 gain scales, one per joint. A single number xi sets both gains of joint i through a sigmoid σ, which squashes it into 0 to 1:

Pi = 3.0 + 4.0 · σ(xi)    Di = 0.03 + 0.02 · σ(xi)  // nominal gains 3.0 and 0.03; P can more than double

Why not output motor torques directly? Earlier work found torque outputs learn locomotion poorly, and that targets plus variable gains suit tasks that need gentle contact. Quiet walking is exactly such a task. The network itself is small: an actor and a critic, each three layers of 128 units, trained with PPO (a standard reinforcement learning algorithm) using legged_gym’s settings.

Animated. As σ sweeps from 0 to 1, P goes from 3.0 to 7.0 and D from 0.03 to 0.05. The spring shows the effect: soft gives way under the same load, stiff holds it (illustrative).
Paper Fig. 2. The whole training setup. aibo’s motion sensor, joint encoders and paw switches feed the policy; it outputs joint targets and a gain scale (stiff vs. damped); many aibos train in parallel in the simulator; and the penalty weights step up between the two curriculum phases.
What the policy sees (Table I)SizeTraining noise
Joint positions120.01 rad
Joint velocities121.5 rad/s
Last joint targets120
Last gain scales120
Paw contact switches40
Gravity direction (body tilt)30.05

What this means: the policy works only from aibo’s own body sensors, 55 numbers in total, and nothing about the room. Knowing its own last gain settings lets it change stiffness smoothly.

Scene 4

Feeling the floor

In short: a cheap on/off switch in each paw tells the policy when the paw lands, so it knows when to soften the leg and when to stiffen it. Without the switches, aibo didn’t learn to walk at all.

Bigger research robots can sense contact with force sensors or estimate it from motor currents. aibo is built to be affordable for an ordinary household and has no force or torque sensors. Estimating contact from a model is not accurate enough on a small, mass-produced robot, because friction and the exact dynamics are hard to pin down. So the policy reads aibo’s paw switches directly. They only say on or off, but that is enough to time the change from soft to stiff.

Fig. 5 shows what the trained policy does with one leg over a single 0.7-second step. In swing (A) the ankle and shoulder-roll gains are high to move the paw forward. Just before touchdown (B) the gains fall, so the leg lands soft. Once the paw is down (C), the shoulder-pitch gain climbs to its maximum to carry the body while the other legs move. The paw’s speed falls just before contact.

Animated. The right foreleg’s three gain scales over one step, redrawn by eye from Fig. 5 (approximate). Watch every gain sink around touchdown (B), then the shoulder pitch rise once the switch is pressed.
Paper Fig. 5. Measured gain scales (top) and paw speed from joint readings (bottom) for the right foreleg; the shaded area is when the paw switch reports contact.

Ablation: no paw switches. Training still converged, but the resulting aibo did not walk. At the start of the quiet-walking phase it fell often, which ends training episodes early. The authors’ explanation is that without contact information the policy cannot tell when to stiffen the joints to hold the body up, so standing still becomes the cheapest way to satisfy the strong noise penalties.

Scene 5

Walk first, then tiptoe

In short: punish noise from day one and the robot learns to stand still. So training has two phases: first learn to walk, then turn up the noise penalties.

The authors’ first attempts without a curriculum failed. Training either did not converge or aibo learned to stay put, a strong “local minimum” (a solution that looks good to the learner but is a dead end) that is perfectly quiet and useless. The fix is a curriculum, inspired by work that splits training into stages to balance exploring and exploiting.

  • Noisy-walking phase. The three noise penalties are small. aibo learns basic walking and following speed commands, even if it stomps.
  • Quiet-walking phase. Starts once the summed velocity-tracking reward of an episode exceeds 1.5. The paw-speed penalty is multiplied by 5; the joint- and body-acceleration penalties are doubled.
Animated. At the switch, the paw-speed penalty grows 5× and the other two 2×. All the other rewards stay the same.
Noise penalty (Table II)Noisy phaseQuiet phase
Foot contact velocity ‖vf‖²−5.0−25.0
Joint acceleration ‖φ̈‖²−2e-7−4e-7
Body angular acceleration−5e-5−1e-4

What this means: the curriculum only reweights three terms. The same reward for following speed commands, and the usual penalties for jerky, slipping or tilting motion, are active throughout.

Curriculum ablation (Fig. 4)SoundMax slope
Proposed: two phases22.7 dB3°
Mild (noisy-phase) penalties only≈30.4 dB5°
Strict (quiet-phase) penalties from the starttraining did not converge; no walking

What this means: mild penalties alone give only about half the noise reduction, and strict penalties alone never produce a walk. Both phases are needed. (≈ values are read from Fig. 4.)

Scene 6

The quietest walk

In short: on the real aibo the learned quiet walk was the quietest of four controllers at every speed tested, including Sony’s hand-tuned quiet mode.

The policy trained in simulation runs on the real aibo without any fine-tuning. Four controllers were compared, all set to walk forward at the same speed:

  • RL baseline: a standard learned gait (legged_gym rewards tuned for aibo), without gain outputs, noise penalties or curriculum.
  • Sony normal and Sony quiet: aibo’s commercial walking controllers, carefully hand-tuned by Sony. They are proprietary, so they were compared as black boxes.
  • RL proposed: this paper’s policy.

Sound was recorded with one of aibo’s own microphones, on the rear side of its head, at 48 kHz. The analysis kept only the range people hear (20 Hz to 20 kHz) and averaged the power spectrum over 85 ms windows (Welch’s method). The microphone is only about 10 cm from the feet, so a person across the room hears much less. Walking gets louder with speed, since the paw has less time to slow down, so the comparison was repeated at several speeds between roughly 5 and 11 cm/s.

Paper Fig. 3. Sound level against measured walking speed. The proposed policy (red) sits below all three others at every speed, including Sony’s quiet controller (green).
Real run. The real aibo walking with the proposed policy, the run the authors’ video labels 22.7 dB. The loop is muted; the project page has the sound. Footage: QuietWalk project page, Sony Group Corporation and RSL, ETH Zurich.
At ≈10 cm/s (Fig. 3/4)Soundvs. proposed
RL baseline35.9 dB+13.2 dB
Sony normal32.9 dB+10.2 dB
Sony quiet≈26.6 dB≈+3.9 dB
RL proposed22.7 dB—

What this means: each 10 dB is a factor of ten in sound power, so 13.2 dB below the standard learned gait is roughly a twentieth of the sound power (our arithmetic). The ≈ value is read from the plot; 35.9, 32.9 and 22.7 dB are also printed in the authors’ result video.

Scene 7

Quiet vs sturdy

In short: the quietest walk handles the smallest slopes. How varied the training world is works as a knob between quiet and sturdy.

To measure robustness, aibo had to cross a 0.5 m ramp within 20 seconds, with the ramp made steeper step by step; the steepest ramp managed is the score. The paper treats coping with unfamiliar ground as one important aspect of robustness. Across all eight controllers, the quieter ones tend to climb less: the loud baseline makes 7°, the quietest policy 3°.

During training, the simulator already varies the world. This is domain randomization: body mass ±0.4 kg, random pushes to the body’s speed every 4 s, small external forces and torques, floor bumps of 0.2 to 1 cm, and floor friction between 0.4 and 0.7. More variety usually makes a gait more robust but louder, a trade-off also reported for learned drone racing. Two variants were tried:

  • More friction variety (0.2 to 0.9): climbs 5° at ≈28.9 dB.
  • Bumpier floor (bumps up to 3 cm instead of 1 cm): climbs 7°, as steep as the baseline, at ≈27.4 dB, the third-quietest result.
Real run. The quietest policy on the 7° ramp: its paws keep stepping, but it makes no headway (its best is 3°). Footage: QuietWalk project page, Sony Group Corporation and RSL, ETH Zurich.
Real run. The loud RL baseline on the same 7° ramp walks straight up and out of the frame. Footage: QuietWalk project page, Sony Group Corporation and RSL, ETH Zurich.
Animated. The only change between the two policies is the size of the random floor bumps during training. Bump heights are drawn exaggerated.
Paper Fig. 4. Each dot is one controller: how loud it is (up) against the steepest slope it climbed (right). The ideal would sit bottom-right. Red dots are the paper’s methods.
Controller (Fig. 4)SoundMax slope
RL baseline35.9 dB7°
Sony normal32.9 dB6°
Fixed PD gains (no gain output)≈31.8 dB7°
No curriculum (mild penalties)≈30.4 dB5°
More DR: friction≈28.9 dB5°
More DR: terrain height≈27.4 dB7°
Sony quiet≈26.6 dB5°
RL proposed22.7 dB3°

What this means: letting the network change its joint stiffness is worth about 9 dB on its own (fixed gains ≈31.8 dB vs 22.7 dB). A bumpier training floor recovers the baseline’s 7° slope while staying about 8.5 dB quieter than it. (≈ values are read from Fig. 4.)

Wrap-up

Limits and what’s next

In short: quiet costs some sure-footedness, the paw is not the only source of noise, and the proxy idea could be used for other things a simulator can’t model.

  • Robustness. The quietest policy is the least robust on slopes. The authors suggest using perception to estimate floor friction, height and softness, and choosing a policy to match.
  • Other noise. Only footstep sound was targeted. Motor and gear noise and mechanical friction also add up, and are left for future work.
  • The general trick. Rather than simulate the sound itself, minimize a simulated quantity that goes with it. The authors point out the same idea could apply elsewhere, for example minimizing torque in simulation to extend battery life.

The authors describe this as the first work to demonstrate the key factors that learned, simulation-trained locomotion needs to walk quietly on a real robot: variable joint stiffness, paw contact switches and a two-phase curriculum.