CoRL 2024 · an animated walkthrough

Learning to Open and Traverse Doors with a Legged Manipulator

The paper in 98 seconds · narrated · sound on
Transcript

Doors are everywhere in our buildings. This robot opens them on its own, push or pull, hinge left or right, even when it gets pulled around. No one tells it which way a door opens. So how does it know?

A walking robot with an arm could reach almost any room, if it could get through doors by itself. But a door hides what it needs. Two doors can look the same, yet one opens by pushing and the other by pulling. Earlier robots had to be told the direction, or ran a fixed push-then-pull test.

The idea: in simulation, a teacher learns by trial and error (reinforcement learning) while it can see hidden facts like the door type. Then a student with a memory (a recurrent network) learns to copy the teacher from real sensors alone. Every practice run brings a new door: push or pull, left or right, heavier, stiffer, wider.

On the real robot, it first guesses pull and backs up. The door doesn’t move. At 2.7 s its guess flips to push, and it drives forward.

It got through a spring-loaded door 38 times out of 40, and opened it every time. Here are those trials, sped up: pushing on one side of the door, pulling on the other. It also recovered when a chair, a person or a rope got in its way, none of which it saw in training. As far as the authors know, it is the first single controller that works out push or pull by itself.

Robot footage: Robotic Systems Lab, ETH Zurich, from the paper’s video. Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

A four-legged robot with an arm learned, in simulation, to open a door and walk through it without being told whether to push or pull. It works that out by feel, like a person at an unfamiliar door.

  1. Why this matters

    Our buildings are made for people, and doors are everywhere in them. A walking robot that can’t get through a door is stuck in one room. One that can open doors on its own can reach most of the building.

  2. What makes it hard

    A door hides what it needs. At a new café you push; if it doesn’t budge, you pull. Doors also differ in spring stiffness, and some swing shut on their own. The robot has to move legs and arm together, decide when to let go of the handle, and hold the door while it squeezes through.

  3. What people did before

    Planning-based controllers need the door’s size and opening direction given in advance. Learning methods coped with different weights and springs, but treated push and pull doors as separate tasks, or found the direction with a fixed push-then-pull test. Most stopped once the door was open.

  4. What this paper does

    A “teacher” learns the task by trial and error in simulation, where it can see hidden facts such as the door type. A “student” with a short-term memory then learns to copy the teacher using only what the real robot can sense, so it has to infer the hidden facts from how the door reacts.

  5. What they showed

    On the real ANYmal robot the student got through a spring-loaded door 38 times out of 40 (95.0 %), and opened it every time. It handled all four door types (push or pull, hinge left or right) and recovered when a chair, a person or a rope got in its way.

  6. Why it’s a step forward

    According to the authors, it is the first single controller that handles both push and pull doors without being told which, or running a fixed test routine. It still needs an external tracker to find the handle, and sometimes doesn’t turn the handle far enough.

Words used below
Policy
the robot’s controller: a neural network that maps what it senses to motions.
Reinforcement learning (RL)
learning by trial and error, guided by a score (the reward).
Teacher–student training
a teacher that sees hidden simulator facts trains a student that sees only real sensors.
Privileged information
facts only the simulator knows, like the door type or spring torque.
Recurrent network
a network with memory: it carries a running summary of recent moments.
Proprioception
the robot’s sense of its own body: joint angles, speeds, body tilt.
1 / 7
simulation-only facts what the robot senses learned policy push pull
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper (CC BY-NC-ND 4.0); the text is a plain-language walkthrough.

Scene 1

Buildings are full of doors

In short: a walking robot with an arm can only use a building made for people if it can open doors and walk through them by itself.

A legged manipulator is a walking robot with an arm. The legs take it over steps, clutter and uneven ground; the arm lets it touch and move things. The authors call opening and passing through doors an essential skill for such robots, because it “significantly” expands where they can go in human-centred places.

Opening a door feels trivial to us, but it is a hard control problem for a machine with this many joints. The robot has to walk up, catch the handle, turn it, swing the door, let go at the right moment, and sometimes hold the door against its closer while it squeezes through a doorway not much wider than its body.

Paper Fig. 1. The same policy on a pull door (top row) and a push door (bottom row). Nobody told it which one it was facing.

The ideal the authors aim for is a “plug-and-play” controller: give it only the position of the door, and it works the rest out while it goes, as a person does at a door they have never seen.

Scene 2

A door hides what it needs from you

In short: the most important fact about a door, whether it opens by pushing or pulling, can’t be seen. Earlier systems had to be given it.

Doors differ in ways you can see (width, handle height) and in ways you can’t: the push or pull direction, the stiffness of the spring that resists opening, the weight of the panel, and whether a closer pulls the door shut. The last one matters for getting through: if the robot lets go too early, the door swings back into it.

Animated. The moment the robot lets go, the closer swings the door shut. To pass through, it has to keep the door open with its arm or body. Illustrative dynamics.

Where earlier approaches stop

ApproachWhat it achievesWhere it stops
Model-based control on wheeled and legged robots (tracking references, model predictive control)Opens doors by following a planned motionNeeds the door’s properties in advance, including its opening direction and precise size
Commercial door behaviour (Spot with arm)A working product featureIts documentation says it must be given the opening and swing direction first
Learning: DoorGym; curiosity-driven RL on a wheeled-legged robotRobust to door mass, spring resistance and frictionPush and pull are separate tasks; they stop once the door is open
Ito et al. (2022), wheeled mobile manipulatorOpens and passes through doorsLearns from human teleoperation demos; separate modules for push and pull
Kang et al. (2023), wheeled mobile manipulatorOpens and passes through push and pull doorsRL is used only for push doors; a pre-programmed push-then-pull test finds the direction

What this means: every earlier system either had to be told the opening direction or found it with a fixed test routine. None had one learned controller that works it out on its own.

Scene 3

A teacher who sees everything, a student who learns to feel

In short: a teacher learns the task with hidden simulator facts, then a student with memory learns to copy it from real-world senses only, and so learns to infer the hidden facts.

Why not learn directly from the robot’s real senses? Because the door type is exactly what can’t be sensed. The same idea worked well for walking (Lee et al. 2020; Miki et al. 2022): there, a student learned to guess the terrain under its feet. For doors, the missing information can only be gained by touching the door, so the approach suits the task even better.

Paper Fig. 2. The teacher (left) sees privileged door facts and learns with RL. The student (right) sees only noisy body and door-position measurements, feeds them through a recurrent memory, and is trained to copy the teacher and to guess the hidden facts.

What each one sees

  • Both: the robot’s proprioception (body orientation, body speed, arm joint positions) and where the handle and the doorway are relative to the robot.
  • Teacher only: the door’s hinge and handle angles and speeds, panel mass, spring and damping torques at hinge and handle, the door type, and which stage of the task it is in. No noise.
  • Student: Gaussian noise on every input. The previous action and the arm joint speeds are left out: training with much more noise on arm joint speed transferred better to the real robot, so the authors removed that input entirely.

The teacher is trained with PPO, a standard RL algorithm. The student is trained to reproduce the teacher’s actions (imitation loss), and a decoder also predicts the door and handle position, the door’s joint states, mass, torques and door type from the student’s memory (estimation loss). The network follows the student of Miki et al. (2022) for walking, but with a plain linear decoder in place of the attention gate: unlike walking, the door-position measurements are assumed noisy but never useless.

Animated. From a single noisy snapshot you can’t tell whether the door is moving; a memory averages over the last moments and can. Illustrative signals.

Is the memory needed? Is the teacher needed?

The authors swapped the student’s recurrent memory for a plain feed-forward network (an MLP) that sees only the current step. It failed to imitate the teacher: with noisy inputs, one snapshot can’t tell apart states that need different actions.

Paper Fig. 9. Imitation loss while training the student: with memory (RNN) it keeps falling; without it (MLP) it stalls at a higher level.

They also trained teachers with RL directly on the student’s inputs (no privileged facts), with and without noise, and tested them on 4,000 random doors in simulation for 10 episodes of 10 seconds each. Without the privileged facts, the policy could not learn to pass through doors of both opening directions. It found an opening trick that sometimes works for both (moving between certain arm poses until the door gives), with a much lower success rate, and noise made that trick worse.

Paper Fig. 13 (appendix). Success rates in simulation. The nominal teacher (blue) opens and passes through both door types; teachers trained only on the student’s inputs (orange, and green with noise) do far worse, and pass through pull doors not at all.

What this means: both halves are needed. The privileged teacher makes the task learnable, and the student’s memory makes it work from real, noisy senses.

Scene 4

Thousands of different doors, in simulation

In short: the policy trains on a simulated ANYmal with a hook, facing a new random door every episode, so that the real door is just one more variation.

Training runs in the Isaac Gym simulator with a model of ANYmal carrying an arm (the same robot used in the experiments). The door has two joints: the hinge (angle θ) and the handle (angle φ). Springs that resist opening the door and turning the handle are modelled as constant torques. A door closer is modelled as damping that grows with the door’s speed, plus a little extra that grows with speed squared for air resistance. Four door types: {pull, push} × {hinge right, hinge left}.

Paper Fig. 3. The training scene. The inset shows the hook the arm uses to catch the lever handle.
Animated (3-D). The hook comes over the lever and presses it down; only then can the door move. Motion is illustrative.

What gets randomized every episode

PropertyRange
Door
Doorway width0.8–1.0 m
Panel thickness0.02–0.06 m
Handle height · length · offset from the panel0.7–1.3 · 0.08–0.12 · 0.03–0.12 m
Panel mass15–75 kg
Hinge spring torque (0 with prob. 0.2)0–30 Nm
Handle spring torque (0 with prob. 0.2)0–3 Nm
Air-resistance coefficient0–4 Nms²
Closer damping = hinge spring × α (all damping off with prob. 0.4)α = 1.5–3 s
Maximum handle turn15–90°
Robot
Start: distance from the wall · sideways from the doorway centre1–2 m · ±2 m
Start heading · start speedany · ±0.5 m/s
Arm joint gains (proportional · damping)40–60 · 3–6

What this means: the policy never sees the same door twice, so it has to learn a way of opening doors that works across all of them, including real ones it has never met.

Paper Fig. 12 (appendix). The randomized door and handle dimensions.

Two choices stand out. Handle friction is set to zero: many real handles are slippery and hard to turn with a hook, and with friction the policy learned aggressive moves that used the hook’s momentum to flick the handle. And the start pose is random, including the heading, so the robot must find the door from anywhere within 2 m.

What the policy controls

The policy steers the arm directly and the legs indirectly. For the legs it sends walking commands (forward, sideways and turning speed) to an existing learned walking controller (Ma et al. 2022), capped at 0.5 m/s and 1 rad/s; commands below 0.1 are set to zero. For the arm it sets joint targets for the arm’s motors. For safety, each target is clipped to a band around the joint’s current angle. The band width comes from the motor’s torque limit and gain, scaled by σ = 0.7, as a stand-in for keeping the motor below its torque threshold (Eq. 1). The same clip runs in training and on the real robot.

targeti = clip( s·ai + q̃i , qi − σ·τ̄i/Kp , qi + σ·τ̄i/Kp )  // a: action, q̃: default pose, q: current angle, τ̄: torque limit, Kp: gain, σ = 0.7
Animated. However far the policy asks the joint to jump, the command sent stays inside a band around where the joint is now. Illustrative signals.
Scene 5

Two stages: open the door, then get through it

In short: the teacher is scored first for opening the door, then for moving through it; pull doors get extra points for getting around the panel.

Rewards are easier to write for each half of the task on its own, so the teacher’s task has two stages.

  • Stage 1, open. Points for moving the hook to the handle, holding a good hook orientation, having the handle inside the hook, and turning it, with a penalty for losing the grip. A separate term rewards opening the door towards a target of 75°. Once the door is past 30°, the handle points are paid in full, so the policy can let go without losing anything.
  • Stage 2, pass. Starts when the door is open more than 70°. The reward is the body’s speed along a progress direction p, divided by the top walking speed and capped at 1. Before the doorway, p points to its centre; after it, p points straight through. The cap stops tricks like flinging the arm to move the body faster than the walking controller allows. Stage 2 also pays the full stage-1 reward, so the policy is never tempted to stay in stage 1.
  • Always on (shaping). Penalties for collisions of body, thighs and arm, for fast or jerky arm motion, for body tilt beyond 8°, for stretching the arm to its singular limit, and for commands outside their limits.

Pull doors were the hard case. The progress reward alone did not teach the robot to get around a panel that swings towards it. So once a pull door is open far enough, the robot also scores for having its body and hook in zone Z1 (+1) or the zone behind the panel Z2 (+2), and it must be behind the panel before stage 2 starts.

Animated. As body and hook move from outside into Z1 and then behind the panel into Z2, each earns +1 and then +2. Zone shapes after Fig. 4.
Paper Fig. 4. The two reward zones for pull doors, defined relative to the door panel.
Scene 6

Try one way; if the door won’t move, change your mind

In short: the student keeps a running guess of the door type, and its actions follow that guess. When pulling doesn’t work, the guess flips and it pushes.

Because the student has a decoder, we can read its current guess of the door type while it works. On the real robot the guess moved between push and pull while the policy was unsure, and the actions moved with it. In the run of Fig. 6, a push door with the hinge on the left, the policy first leaned towards pull and drove its body backwards. The door didn’t follow. At 2.7 s the guess switched to push and the policy drove forward. In the authors’ words, it “learned to push and pull until it could move the door panel”.

Paper Fig. 6. Top: the policy’s guess for each door type. Bottom: its forward speed command. The dashed line at 2.7 s marks the flip from pull to push, and the command turns from backwards to forwards.
Real run, half speed. The bars at top left are the policy’s live guess of the door type, read from its decoder. It first favours pull left, then switches to push left and the robot pushes the door open. Footage: Robotic Systems Lab, ETH Zurich, from the paper’s video.

Where does the guess live? The authors projected the student’s memory (its hidden state) to 2-D with UMAP, a method that draws high-dimensional points on a flat map. Runs on pull and push doors start in the same central cluster and then head off in separate directions, and the two regions look linearly separable.

Animated. Each run starts undecided in the middle and drifts towards its door type as the door responds. A sketch, not the real data.
Paper Fig. 7. The real projection, from simulated runs on pull-right (red) and push-right (blue) doors. Colour gets darker as the run progresses.

Does guessing the door type help learning?

Surprisingly, no. A student trained without the door-type estimation loss still learned to tell push from pull, and its imitation loss trained just as well. The authors kept the decoder anyway, because it lets them see what the policy believes.

Paper Fig. 8. Imitation loss with (blue) and without (orange) the door-type estimation loss: almost the same.

What this means: the push-or-pull skill comes from copying a teacher who knew the answer. The memory learns to figure it out because that is the only way to act like the teacher.

Scene 7

On the real robot: 38 of 40 doors passed

In short: trained only in simulation, the policy passed a real spring-loaded door 95 % of the time, handled all four door types, and shrugged off disturbances it never trained for.

The student went straight onto the real ANYmal with no fine-tuning, running at 50 Hz, the same rate as in training. Its body senses come from the robot’s own sensors, and the robot tracks its body pose with onboard lidar odometry. The handle and doorway positions come from an external tracker: either a motion-capture system (accurate, but it needs fixed cameras) or AprilTags, printed markers on the door watched by an external camera (easy to put on any door).

Repeatability

Spring-loaded door, motion capturePassed throughOpened
Pull side, 20 back-to-back trials20 / 2020 / 20
Push side, 20 back-to-back trials18 / 2020 / 20
Overall95.0 %100 %

What this means: the robot opened the door every time; the two misses on the push side happened afterwards, when it got stuck on the side of the doorway, which stuck out because the test door had no walls around it, unlike the simulated doors.

Real trials, sped up. The same spring-loaded door from its push side (left) and its pull side (right); the counters end at 18/20 and 20/20. Footage: Robotic Systems Lab, ETH Zurich, from the paper’s video.

All four door types

Doors of every opening and swing direction, and with different doorway and handle sizes, panel inertia, handle dynamics, with and without a closer: the policy worked out the type during the task and passed through each.

Paper Fig. 5 (a–d). Pull right, pull left, push right, push left. Under each photo, the policy’s guess for each door type over time; the solid line is the true type.

Disturbances it never trained for

A chair blocking the door, a person pushing on the door, a rope pulling the robot: none of these were in training, yet the policy recovered. The authors’ explanation: the student never copies the teacher perfectly, so during training its own mistakes put it in odd situations, and the teacher showed how to recover each time. If the student missed the handle, the teacher’s next actions were to try again.

Paper Fig. 10 (a–c). Chair blocking the door, a person pushing on the door, a rope pulling on the robot.

How stiff a door can it handle?

In simulation the authors raised the hinge spring beyond anything in training (0–30 Nm). Opening and passing both get harder as the spring stiffens. Past 50 Nm the robot can still sometimes open the door, but getting through drops to near zero: a stiff push door is hard to hold open and can trap the robot in the doorway, and a stiff pull door closes before the robot gets around it.

Paper Fig. 14 (appendix). Success rate against hinge spring torque in simulation. Shaded: stiffer than anything seen in training.
Wrap-up

Limits and what’s next

In short: the policy works but still leans on an outside tracker, and it can be fooled when it can’t tell whether the handle actually turned.

  • Not turning the handle far enough was a common failure. The policy then gets stuck rocking back and forth, pushing and pulling as its guess of the door type swings. It has no direct sense that tells it the handle turned, so it can confuse its own turning motion with a turned handle. Vision or force sensing could fix this.
  • AprilTags hurt performance because their noise and delay were not modelled in training. With delayed handle positions the policy sometimes let go of the handle and grabbed for it again.
  • Door shapes not in the simulator caused failures: the hook catching on the door panel, or the body catching on a doorway that sticks out.
  • Next: measure the door with the robot’s own sensors instead of an external tracker, add force sensing, and handle more kinds of handles, such as knobs.

The bigger picture: by copying a teacher that knew the answers, a single policy learned to find out what a door needs by interacting with it. That is the step from “tell the robot what kind of door this is” to “let it find out”.