DAPPER trains a fresh robot controller each round and asks people only questions they can actually answer — so a walking robot learns the style someone wants from far fewer answers.
Why this matters
Four-legged robots can learn to walk by trial and error in simulation. But people want a particular style — how high the body rides, whether it stays level, the rhythm of the legs. Getting that today takes an expert who tunes a reward formula by hand, or a skilled operator who can demonstrate the gait.
What makes it hard
Anyone can watch two videos and say which is closer to what they want. But that only teaches the robot something if one clip is clearly closer. When both are equally good or equally off, the honest answer is “can’t decide” — like an eye test where lens one and lens two look the same. In the authors’ study, about 60 % of the hardest-to-tell pairs got exactly that answer.
What people did before
Learning from such answers (preference-based learning) usually shows two clips of the same controller, which is updated a little each round, so its clips look alike. Earlier tricks save questions by auto-labelling confident pairs, reusing old data or picking uncertain pairs — but none of them model how hard a question is for the person answering it.
What this paper does
Two ideas. First, after every batch of answers, train a brand-new controller from scratch and compare whole controllers against each other, so the options genuinely differ. Second, learn from past “can’t decide” answers which pairs a person can tell apart, then ask about those and reward new controllers for staying distinguishable. The authors call this DAPPER.
What they showed
On the ANYmal robot in simulation, DAPPER reached the target posture with about a tenth of the questions earlier methods needed; when questions were hardest to answer, the earlier methods did not get there within 2,000. With GPT-5 answering instead of a person, label accuracy rose from 62 % to 85 %. A crawl learned in simulation ran on the real robot.
Why it’s a step forward
It treats “is this an easy question for a person?” as part of the learning problem, which is what makes preference learning practical for non-experts. Honest limits: the style is described by a few hand-picked features (up to six), it still takes hundreds of answers, retraining from scratch costs compute, and the large comparisons used a simulated answerer.
- Policy
- the robot’s controller — its “brain” that turns sensor readings into joint commands.
- Query
- one question: two walking videos, which is closer to what you want?
- Preference discriminability
- how easily a person can tell which of two clips is closer to their ideal.
- Policy bias
- successive versions of one policy stay alike, so their clips look alike.
- Reward model
- a network that learns, from answers, how much a person likes a motion.
- Discriminator
- a network that predicts whether a person could tell two policies apart.