A four-legged robot with an arm learned, in simulation, to open a door and walk through it without being told whether to push or pull. It works that out by feel, like a person at an unfamiliar door.
Why this matters
Our buildings are made for people, and doors are everywhere in them. A walking robot that can’t get through a door is stuck in one room. One that can open doors on its own can reach most of the building.
What makes it hard
A door hides what it needs. At a new café you push; if it doesn’t budge, you pull. Doors also differ in spring stiffness, and some swing shut on their own. The robot has to move legs and arm together, decide when to let go of the handle, and hold the door while it squeezes through.
What people did before
Planning-based controllers need the door’s size and opening direction given in advance. Learning methods coped with different weights and springs, but treated push and pull doors as separate tasks, or found the direction with a fixed push-then-pull test. Most stopped once the door was open.
What this paper does
A “teacher” learns the task by trial and error in simulation, where it can see hidden facts such as the door type. A “student” with a short-term memory then learns to copy the teacher using only what the real robot can sense, so it has to infer the hidden facts from how the door reacts.
What they showed
On the real ANYmal robot the student got through a spring-loaded door 38 times out of 40 (95.0 %), and opened it every time. It handled all four door types (push or pull, hinge left or right) and recovered when a chair, a person or a rope got in its way.
Why it’s a step forward
According to the authors, it is the first single controller that handles both push and pull doors without being told which, or running a fixed test routine. It still needs an external tracker to find the handle, and sometimes doesn’t turn the handle far enough.
- Policy
- the robot’s controller: a neural network that maps what it senses to motions.
- Reinforcement learning (RL)
- learning by trial and error, guided by a score (the reward).
- Teacher–student training
- a teacher that sees hidden simulator facts trains a student that sees only real sensors.
- Privileged information
- facts only the simulator knows, like the door type or spring torque.
- Recurrent network
- a network with memory: it carries a running summary of recent moments.
- Proprioception
- the robot’s sense of its own body: joint angles, speeds, body tilt.