IEEE Robotics & Automation Magazine · 2026 · an animated walkthrough

DAPPER: Discriminability-Aware Policy-to-Policy Preference-Based Reinforcement Learning for Query-Efficient Robot Skill Acquisition

The paper in 85 seconds · narrated · sound on
Transcript

A walking robot can learn many styles of moving. This low crawl was learned from answers to one simple question: which video is better?

Normally, choosing a robot’s walking style takes an expert, who tunes rewards by hand. But anyone can pick the better of two videos. The catch: both videos usually come from one robot brain, a policy, that changes only a little. They look alike, and people answer: can’t decide.

DAPPER trains a brand-new policy every round, and compares whole policies. Fresh ones spread out, so the choices differ clearly. What counts is preference discriminability: can you tell which is closer to what you want? Very different robots can still be a toss-up. So a second network, a discriminator, learns which pairs people can tell apart. DAPPER asks about those, and rewards policies that stay distinct.

In simulation, DAPPER needed about a tenth of the questions. On the hardest setting, older methods failed within two thousand. With GPT-5 as the judge, can’t-decide answers halved. And a crawl learned this way ran on the real robot, with no extra tuning. Ask questions people can answer, and every answer counts.

Footage: NAIST Robot Learning Lab, from the authors’ supplementary video (real-robot crawl on ANYmal; simulated gaits). In the paper’s experiments the answers come from a simulated annotator calibrated on a study with real people. Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

DAPPER trains a fresh robot controller each round and asks people only questions they can actually answer — so a walking robot learns the style someone wants from far fewer answers.

  1. Why this matters

    Four-legged robots can learn to walk by trial and error in simulation. But people want a particular style — how high the body rides, whether it stays level, the rhythm of the legs. Getting that today takes an expert who tunes a reward formula by hand, or a skilled operator who can demonstrate the gait.

  2. What makes it hard

    Anyone can watch two videos and say which is closer to what they want. But that only teaches the robot something if one clip is clearly closer. When both are equally good or equally off, the honest answer is “can’t decide” — like an eye test where lens one and lens two look the same. In the authors’ study, about 60 % of the hardest-to-tell pairs got exactly that answer.

  3. What people did before

    Learning from such answers (preference-based learning) usually shows two clips of the same controller, which is updated a little each round, so its clips look alike. Earlier tricks save questions by auto-labelling confident pairs, reusing old data or picking uncertain pairs — but none of them model how hard a question is for the person answering it.

  4. What this paper does

    Two ideas. First, after every batch of answers, train a brand-new controller from scratch and compare whole controllers against each other, so the options genuinely differ. Second, learn from past “can’t decide” answers which pairs a person can tell apart, then ask about those and reward new controllers for staying distinguishable. The authors call this DAPPER.

  5. What they showed

    On the ANYmal robot in simulation, DAPPER reached the target posture with about a tenth of the questions earlier methods needed; when questions were hardest to answer, the earlier methods did not get there within 2,000. With GPT-5 answering instead of a person, label accuracy rose from 62 % to 85 %. A crawl learned in simulation ran on the real robot.

  6. Why it’s a step forward

    It treats “is this an easy question for a person?” as part of the learning problem, which is what makes preference learning practical for non-experts. Honest limits: the style is described by a few hand-picked features (up to six), it still takes hundreds of answers, retraining from scratch costs compute, and the large comparisons used a simulated answerer.

Words used below
Policy
the robot’s controller — its “brain” that turns sensor readings into joint commands.
Query
one question: two walking videos, which is closer to what you want?
Preference discriminability
how easily a person can tell which of two clips is closer to their ideal.
Policy bias
successive versions of one policy stay alike, so their clips look alike.
Reward model
a network that learns, from answers, how much a person likes a motion.
Discriminator
a network that predicts whether a person could tell two policies apart.
1 / 7
one policy (earlier methods) the person’s answers DAPPER target style
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. Figures are from the paper; the text is a plain-language walkthrough.

Scene 1

Why it matters

In short: robots can learn many ways of walking, but today only experts can tell them which way — answering “which video is better?” could open that up to anyone.

The paper starts from a simple wish: that a legged robot walk the way its user prefers, not just the way a task formula dictates. There are two standard routes, and both are expensive for ordinary users. Reinforcement learning — learning by trial and error to maximise a reward — needs a reward function with several hand-tuned terms, which takes expert knowledge. Imitation learning needs demonstrations, and demonstrating a four-legged gait through a robot interface takes skill and a clear idea of the target motion in advance.

A third route asks much less of the user: show two videos and ask which is closer to what they want. This is preference-based reinforcement learning (PbRL). From the answers, a reward model learns to score motions the way the person would; the robot's controller is then trained on that score, new videos are shown, and the loop repeats. The authors say this is, to their knowledge, the first application of PbRL to quadruped locomotion. The obstacle is the number of questions: PbRL usually needs so many that it is impractical on a real schedule.

Paper Fig. 1. The idea in one picture: many differently-behaving controllers (π₁ … π₆, real ANYmal photos) are collected, and questions are built from pairs a person finds easy to judge.
Scene 2

The problem: look-alike questions

In short: when both videos come from the same slowly-changing controller, they often look equally good or equally bad, and the person can only say “can’t decide” — a wasted question.

In a typical PbRL setup, both clips of a question are cut from different runs of one policy. After each batch of answers that policy is updated a little — it is a continuation of the previous one. The paper calls the resulting lack of variety policy bias: successive versions behave alike, the clips they produce look alike, and the updated reward is hard to exploit because the policy barely explores anything new.

The key observation is about the person. A question is easy when one clip is clearly closer to what they want. When both are similarly off, or similarly good, people struggle — “not because the trajectories themselves are indistinguishable, but because their underlying preferences are.” Those questions produce no usable label. The paper names the property that matters preference discriminability, and builds the method around making it high.

Earlier ways to save questions each help, but none looks at this: automatically labelling pairs the reward model is already sure about (SURF), exploring before asking and re-scoring old data after each update (PEBBLE), asking about pairs near the reward model's decision boundary (uncertainty sampling), making learning robust to wrong labels, or choosing questions from a fixed set of pre-recorded clips (active reward learning — which cannot create new, easier-to-compare behaviour).

Scene 3

Compare policies, not clips

In short: DAPPER throws away the “keep updating one controller” habit: after each batch of answers it trains a new controller from scratch, and questions compare two different controllers.

Training from scratch means the new policy is not anchored to the old one, so policy bias disappears and behaviours spread out. Each finished policy goes into a library Π. A question then pairs the newest policy with one earlier policy from the library (“policy-to-policy”), instead of two episodes of the same policy. The preference model is rewritten accordingly: the probability that policy i is preferred over j is a softmax of the reward model's score of a trajectory from each.

P(πᵢ ≻ πⱼ) = exp RH(G(πᵢ)) / [ exp RH(G(πᵢ)) + exp RH(G(πⱼ)) ]  // G(π): a trajectory rolled out by π; RH: the learned reward model

But variety alone is not enough — two very different robots can still be equally far from what the person wants. That is why the second half of the method, below, estimates discriminability directly.

The paper's Fig. 4(b) shows where the learned policies landed in the “posture” feature space (body height × incline). The Baseline starts low and tilted and creeps along a diagonal toward the target. DAPPER without its discriminability bonus jumps around at first — fresh policies help — but then gets stuck far from the target, at the top of the height range. Full DAPPER scatters widely early on and then collects around the target. The scene above is an illustrative version of that pattern.

Paper Fig. 4. (a) How often questions were answerable, how many questions each method needed, and the best style error over time, at three answerability thresholds. (b) Where each method's policies landed in the height × incline space (colour = question count); the target is at the red cross-hair.
Scene 4

What people can tell apart

In short: the authors measured with real people how big the difference must be before a comparison is answerable, and turned that into a rule their experiments could apply automatically.

Running thousands of questions past real people for every comparison would be impractical, so the experiments use a simulated annotator: a rule that answers like a person would. To make that rule honest, the authors first ran a study. They trained 25 policies on a grid of body heights (29, 37, 45, 53, 61 cm) and incline angles (−8, −4, 0, 4, 8°). Five participants saw a reference video of the target posture — 45 cm high, level body — and two candidate videos, and answered “Left is better”, “Right is better” or “Can't decide”, for the same 105 pairs in random order.

The simulated annotator measures each clip's distance from the target as the sum of its feature errors (each feature scaled to 0–1, divided by √(number of features)). The discriminability of a question is the gap between the two clips' distances. Below a threshold, the answer is “can't decide”; above it, the closer clip wins.

dpref(τ) = (1/√|f|) Σₖ |fₖ(τ) − f*ₖ|    ddisc = | dpref(τᵢ) − dpref(τⱼ) |  // f: features, f*: target
Animated. The same three pairs as the scene. Watch the middle one: A and B are far apart on the map, but equally far from the star, so the gap is zero and the answer is “can't decide”.

What people did: at the smallest gaps, around 60 % of answers were “can't decide”; the share fell as the gap grew and reached zero above 0.4 — every participant could tell. When people did pick a side, they were right more than 80 % of the time at every level, and always right above a gap of 0.3. From this the authors set three thresholds for the experiments: Small = 0.2, Medium = 0.3, Large = 0.4 (larger means harder questions, since more pairs count as unanswerable). One answer took about 7 s on average.

Paper Fig. 3. (a) The 25 study policies on a grid of body heights and incline angles; the target is in the centre. (b) The console participants used. (c) “Can't decide” rate and accuracy against the gap in distance to the target.
Threshold nameGap below which the answer is “can't decide”Matches
Small0.2a sharp-eyed person
Medium0.3the middle of the study's range
Large0.4the gap above which all five participants could always tell

What this means: the experiments that follow simulate people as the study found them — sometimes unable to decide — rather than as perfect judges.

Scene 5

The discriminator

In short: DAPPER learns, from past answers, which pairs a person can tell apart — and uses that both to choose questions and to steer the next controller toward being easy to compare.

Every answer is stored with its label: 1 or 0 when the person chose, 0.5 for “can't decide”. The chosen answers train the reward model RH in the usual PbRL way. All answers — sorted into “could tell” and “couldn't” — train the discriminator D, a network that takes trajectories of two policies and outputs the probability that a person could tell them apart. It is symmetric (the order of the pair doesn't matter), uses Monte Carlo dropout to stay stable early on when it has seen few labels, and is re-trained from scratch every round so it doesn't overfit to the first policies.

Use 1 — choosing the question

One side of every question is the newest policy. Its partner is drawn from the library with a probability that grows with D's score for the pair — a softmax with a temperature α. A higher α picks tellable pairs more often; too high, and it keeps asking about the same easy pairs, which starves the reward model of variety. Pairs already asked are never repeated.

pₓ ∝ exp[ α · D(G(πₓ), G(πᵢ)) ]  // πᵢ: newest policy, πₓ: candidate partner from the library
Animated. As α rises, the choice of partner collapses onto the pairs D rates easiest. Scores and the α range are illustrative; the paper uses α = 10⁻³ in all experiments.

Use 2 — a bonus for being distinguishable

While the next policy trains, its reward mixes the person's preferences with D's opinion: at every step, RD is D's average score for “the trajectory so far” against each earlier policy. So the new controller is pushed to be both liked and easy to tell apart from what came before.

r = (1 − β)·RH + β·RD,   β = 0.6  // RH: preference reward, RD: discriminability reward
Animated. Sliding β: too little bonus and new policies can end up like old ones; too much and exploration crowds out the target. The spread of dots is illustrative; the verdicts follow the paper's Fig. 7.

Why not just reward distance from earlier policies? The paper tests exactly that (“Ours-DM”, which maximises feature distance). It explores more but converges worse — being different isn't the same as being easy to judge.

Paper Fig. 7. Effect of the sampling temperature α (left) and the bonus weight β (right) at the three thresholds. α = 10⁻³ and β = 0.6 worked well across all of them.
Scene 6

The loop

In short: each round trains a fresh controller, asks ten questions and updates two small models; training dominates the time, but far fewer rounds are needed.

The three steps of Fig. 2 repeat until the learned style is close enough to the target (a distance below 0.02):

  1. Policy learning. A new policy starts from scratch and trains on the mixed reward r: 300 updates, each from 128 simulated robots walking 128 steps in parallel. Training uses Lagrangian PPO, a reinforcement-learning method that also enforces constraints — here walking forward at 1 m/s (± 0.2), no sideways drift or turning, no body or leg collisions — so every candidate is a sensible walk and the person only judges style.
  2. Query collection. Ten questions pair the new policy with sampled partners from the library; answers go into the question dataset.
  3. Reward update. The reward model and the discriminator are re-fitted on all answers so far.
Paper Fig. 2. The DAPPER loop: train the current policy on preference and discriminability rewards; sample a partner from the policy library and ask; update the reward model and discriminator from the question dataset.

Retraining from scratch every round sounds wasteful, so the authors timed it (Fig. 6), on one RTX 4090 GPU and assuming about 7 s per human answer. A round takes 13–14 minutes, 12–13 of them policy training; ten questions take about 70 s; updating the models takes seconds. The Baseline can keep training one policy with fewer updates per round. With 30 updates it cut its own total time to about 40 % of the 300-update version — yet still took about twice as long as DAPPER overall, because it needed many more rounds. With only 10 updates per round, asking questions dominated and total time grew again.

Paper Fig. 6. Left: time per round, split into training, asking and updating. Right: total time until the target is reached, DAPPER against the Baseline with 300, 30 and 10 updates per round.
Scene 7

Results

In short: DAPPER needed the fewest questions in every setting, kept working when questions were hardest to answer, made an AI judge far more useful, and produced gaits that ran on the real robot.

All experiments use the ANYmal quadruped in the legged_gym / Isaac Gym simulator; the controller runs at 50 Hz and sends joint targets to 400 Hz joint controllers. DAPPER is compared with the Baseline (classic single-policy PbRL), SURF (auto-labelling) and PEBBLE (pre-exploration and re-scoring), plus two variants of itself: without the discriminability bonus, and with a plain distance bonus (Ours-DM). Each setting was run five times.

Questions needed (Fig. 4)

On the posture task, DAPPER reached the target with the fewest questions and had the highest share of answerable questions at every threshold. At the hardest threshold (Large), only DAPPER and Ours-DM kept producing enough answerable questions: Ours-DM needed more than three times as many questions as DAPPER, and the single-policy methods did not reach the target within the 2,000-question limit. In the Discussion the authors sum up that DAPPER cuts the question count to about one tenth of what previous approaches need — though that is still hundreds of answers.

Across a fine sweep of thresholds (Fig. 5), DAPPER beat the Baseline everywhere. The gap narrows when questions are easy — but even with a threshold of zero, where every question is answerable, the Baseline lags, because policy bias stops it from reaching new behaviour in few rounds. At very high thresholds answerable questions become so rare that DAPPER fails too.

Paper Fig. 5. Questions needed as the answerability threshold rises, DAPPER against the Baseline (capped at 2,000).

More features (Fig. 9)

Two harder styles add gait timing: “Trot only” (4 features: one leg's cycle time and the timing offsets between legs) and “Normal” (6 features: those plus height and incline). DAPPER converged faster in both; earlier methods found very few answerable questions with six features. The authors note even DAPPER may hit limits as features grow.

Paper Fig. 9. Discrimination rate, questions needed and style error for the 4-feature and 6-feature gait tasks (Medium threshold).

An AI as the judge (Table IV)

Could a vision-language model answer instead of a person? The authors gave GPT-5 three frames (at 0.5, 1.5 and 2.5 s) of a reference video and two candidates. Such models rarely admit “can't decide”, so answers where the frame-wise preferences were not significantly different (Welch's t-test) were mapped to “can't decide”. With up to 1,000 questions:

Animated. Well-separated frame scores give a real answer; overlapping ones become “can't decide”. Scores are illustrative.
GPT-5 labels, mean ± std of 5 runsDAPPERBaseline
“Can't decide” rate23.6 ± 15.7 %46.9 ± 16.0 %
Label accuracy84.7 ± 12.8 %61.9 ± 15.2 %
Smallest style error reached (target < 0.02)0.005 ± 0.0020.318 ± 0.061

What this means: with the Baseline, the AI couldn't answer half the questions and was barely better than a coin flip on the rest, so almost nothing was learned; DAPPER's easier questions made the same AI useful enough to reach the target style — even with about 15 % of its labels wrong.

Paper Fig. 8. What the vision-language model was shown: instructions plus sampled frames of the reference and the two candidates.

On the real robot

A “crawl” style (body height 30 cm, level body, trot timing) was learned in simulation and run on the real ANYmal using the standard sim-to-real recipe of legged_gym — randomising sensor noise and physical parameters in training — without extra tuning. The style features are only needed in simulation, to compute the reward.

Real run. The crawl policy learned in simulation, running on the real ANYmal. Footage: NAIST Robot Learning Lab (DAPPER supplementary video).
Paper Fig. 10. Frames of the real-robot crawl over time.
Wrap-up

Limits and next steps

In short: DAPPER makes preference learning much cheaper in questions, but the style is still described by a handful of hand-chosen features, and real users in the loop remain the next step.

  • Hand-picked features. Styles are described by up to six chosen features (heights, angles, gait timings). Conflicting or noisy features could stall learning; learning the features automatically from sensor data is a stated next step.
  • Still hundreds of questions. About a tenth of before is a big cut, but richer, higher-dimensional preferences (smoothness, energy use, style) will need more help. With more features, people may also find more pairs undecidable — and Fig. 5 shows that very high thresholds eventually break the method.
  • Compute for retraining. Training from scratch each round is the price of avoiding policy bias; warm starts or meta-learning could lower it.
  • Early noise in the discriminator. With few labels its estimates can mislead exploration; Monte Carlo dropout helped here, but larger problems may need more.
  • Simulated answers. The large comparisons used the simulated annotator calibrated on five people; the GPT-5 study is a proof of concept. Wrong labels (not just “can't decide”) were tolerated at about 15 %, and combining DAPPER with noise-robust methods is suggested.
  • Beyond walking. Only quadruped locomotion was tested; manipulation or humanoids are possible future targets.