JSAI Annual Conference · 2015 · an animated walkthrough

A drone that sticks its own navigation markers on the ceiling

The paper in 106 seconds · narrated · sound on
Transcript

A small drone, flying on its own. It climbs to the ceiling, and sticks a marker up there. Then it flies on to the next spot, and sticks another. Why would a robot put up its own signs?

Hovering drones could do many jobs indoors. But first, a drone has to know where it is. Its camera reads printed squares called AR markers and measures its position from them. One step past the last marker, and it is lost. Earlier systems used markers laid out in advance, so a drone could only work where people had prepared the room.

This paper’s idea: let the drone put up its own signs. It carries a stack of markers and sticks them on the ceiling, each one next to a marker it already knows. When a known and a new marker share one camera picture, the new one gets its place on the map. The map grows like a chain. With no marker in view, a Kalman filter, a way to blend a drifting guess with exact fixes, keeps the position going on the drone’s own speed. Where should the next marker go? Reinforcement learning, trial and error in a simulator, favoured spots toward the edge, even when markers land off target or fall.

Here is the real drone. It holds itself just under the ceiling and presses a marker into place. On the real drone it worked: five markers stuck in a row, fully automatically. Some markers bunched at one spot, and the learner needs a smaller problem. But the robot built the navigation aids it needed, by itself.

Footage: the authors’ own video (YouTube, Takahiro Miki, The University of Tokyo), cropped to hide its burned-in labels. Animations are illustrative. Voice: Kokoro TTS (synthetic). Music and sound effects: synthesized for this video.

The story in plain words

A small drone that sticks its own navigation markers onto the ceiling, so it can find its way into places nobody prepared for it.

  1. Why this matters

    Small drones that can hover in place are being tried for all sorts of jobs: stacking blocks into structures, or guiding people. Whatever the job, the drone first has to know where it is, or it cannot fly to the right spot.

  2. What makes it hard

    A camera can read printed square markers and work out its exact position from them. But that only works where someone has put markers up. It is like a hiker who can only walk where trail signs are posted: one step past the last sign and they are lost.

  3. What people did before

    Earlier work (Rudol et al., 2010) did locate a small drone from markers, but those markers were laid out on the ground in advance. That works well, as long as a person has prepared the space first, and the drone stays fenced in by it.

  4. What this paper does

    Let the drone put up its own signs. It carries a stack of markers and sticks them on the ceiling, each one next to a marker it already knows, so the known area grows like a chain. A filter keeps its position when no marker is in view, and a trial-and-error learner picks where the next marker should go.

  5. What they showed

    A real, off-the-shelf AR.Drone with its camera pointing up stuck 5 markers in a row, fully automatically. It mapped markers one after another and kept its position when markers dropped out of view. Some markers bunched up at the same spot, and the learner barely improved in its 100 training steps.

  6. Why it's a step forward

    Instead of waiting for people to prepare the room, the robot builds the navigation aids it needs. The next steps the authors name are a smaller learning problem and placing markers where the camera can keep them in view longer, so they land more precisely.

Words used below
AR marker
a printed black-and-white square; a camera can measure its position and angle from it
Self-localization
a robot working out where it is and which way it faces
Kalman filter
blends a fast but drifting guess with slow but exact fixes
Reinforcement learning
learning by trial and error, guided by rewards
ROS
Robot Operating System, software for connecting robot programs
1 / 6
markers drone and its position estimate measuring and deciding
Scene 1

Read the full section with the paper’s figures ↓

The paper, section by section

Everything the animation skips

Each section matches one scene above. Press “Watch scene” to jump back to its animation; click any figure to enlarge it. The original paper is a four-page conference paper written in Japanese; figures are from it, and their Japanese labels are explained in the captions. The text here is a plain-language walkthrough in English.

Scene 1

Why it matters

In short: a drone that navigates by markers can only go where someone has already put markers up. This paper removes that fence.

Quadcopters, drones with four rotors, can hold still in the air, so people expect them to take on many kinds of missions. The paper names two examples: drones that stack blocks into architectural structures (Willmann et al., 2012) and a guide drone (MIT’s SkyCall project). Every such mission needs self-localization: the drone must know where it is.

The paper looks at one way to do this: image markers. An AR marker is a printed black-and-white square. When a camera sees one, software can measure where the marker is relative to the camera and how it is turned. Turn that around and you know where the camera, and the drone carrying it, is.

Earlier work on marker-based localization for small drones (Rudol, Wzorek and Doherty, ICRA 2010) used markers placed on the ground in advance. The limitation the paper points to is simple: the drone can only move within the area where markers have been installed. Someone has to prepare the space before the drone can work in it.

Paper Fig. 9. Markers stuck on the lab ceiling for a first test. The red box (labelled “AR.Drone camera view”) is roughly what the drone’s upward camera sees at once; the arrow points to the reference marker (基準のマーカ) that defines the map’s origin.

So the goal of the paper is a flying robot that, on its own judgement, sticks markers onto the ceiling and so widens the area it can act in.

Scene 2

The idea and the robot

In short: an off-the-shelf drone with its camera turned upward, a marker stack on its back, and a laptop that does the thinking over WiFi.

The system has four parts: a PC, a Parrot AR.Drone, a marker-attaching device, and the markers themselves. The PC talks to the drone over WiFi and steers it by sending velocity commands. The drone’s camera was turned to face up, so it films the ceiling, and the attaching device sits on top of the drone.

The drone

A consumer AR.Drone. Its camera looks at the ceiling; the PC sends it velocities and reads back its measured speed.

The attacher

A holder on top of the drone carrying a stack of markers. The drone flies up and presses the top one against the ceiling.

The markers

Printed squares on card with sticky tabs, stacked alternately so each one does not stick to the one below it.

Real run. The AR.Drone holds itself just under the ceiling to stick the top marker from its stack; markers already on the ceiling are visible around it. Muted loop, cropped to hide the video’s labels. Footage: Takahiro Miki (author’s video), The University of Tokyo.
Paper Fig. 4. The AR.Drone used in the experiments, with the marker-attaching device mounted on top.
Paper Fig. 5. The marker-attaching device: a small holder carrying the stack of markers.
Paper Fig. 6. The front of one marker: the printed pattern on white card, with tabs around it.
Paper Fig. 7. The back of a marker, with sticky patches on the tabs. Markers are stacked alternately so that each one does not stick to the one beneath it.

The software: three programs that talk to each other

The software runs on ROS (Robot Operating System), which lets separate programs, called nodes, run in parallel and pass messages. Three nodes come from the community: ardrone_autonomy (the drone driver), ar_track_alvar (marker recognition) and joy (a gamepad). The authors wrote the other three:

  • marker_sub turns what the camera sees into the drone’s absolute position, and registers the ceiling position of each new marker one by one. It sends the marker map to the brain.
  • send_cmd fuses marker positions with the drone’s measured speed in a Kalman filter, and carries out actions: stick a marker, or fly back to the start point to recharge. A person can switch to manual control at any moment, for safety during experiments.
  • brain combines the marker map, the drone’s position and the battery level into the current situation and chooses the next action with reinforcement learning.
Paper Fig. 8. How the nodes are wired. Camera images go from the drone driver to the marker tracker; marker poses go to marker_sub; its position estimate goes to send_cmd and its marker map to brain; the brain sends the next action back to send_cmd, which commands the drone. Labels are in Japanese.
Paper Fig. 14. The drone during a mission, seen from below, close to the ceiling.
Scene 3

Chaining markers into a map

In short: whenever a known and a new marker appear in the same camera picture, the new one inherits a place on the map. Any known marker then tells the drone where it is.

The marker tracker reports, for each marker in the picture, where it sits relative to the camera and how it is turned. That alone says nothing about the room. The trick is to start from one reference marker whose position is declared known, and grow the map outward from it.

Registering a new marker

Suppose the camera sees one known marker (registered at position xa, ya and angle θa) and one new marker. From the picture, the drone measures the distance l between their centres, the angle α of the line joining them, and each marker’s own angle in the picture (θ0 for the known one, θ for the new one). Then the new marker’s place on the map is:

xm = xa + l·cos(θa + α + θ0)  // paper eq. 2
ym = ya + l·sin(θa + α + θ0)  // eq. 3
θm = θ0 − θ + θa           // eq. 4: the new marker's angle on the map
Animated. Watch the line l and the angles appear; once they are measured, the grey “?” marker becomes a registered one and can anchor the next.

Where am I? From one known marker

Once a marker is on the map, seeing it again is enough to place the drone. If the marker appears at (x, y) in the camera frame with angle θ, then l = √(x² + y²) and α = arctan(x/y), and:

β = θ + θa − α − 180°  // eq. 5
x_drone = xa + l·cos β,  y_drone = ya + l·sin β  // eqs. 6–7
heading = θ + θa  // eq. 8
Animated. As the drone moves and turns, the distance l and angle α to the marker change, and together with the marker’s registered pose they give the drone’s position and heading.
Paper Fig. 10. The first real test, flying under the markers of Fig. 9. Right: the marker positions the drone registered, one after another. Left: its own estimated path (blue) among them. The paper reports that the markers’ positions and angles were recognized.
Scene 4

Filling the gaps

In short: between marker sightings, the drone keeps track of itself from its own speed, and each new sighting corrects the drift.

Marker positions are accurate but come in only about 10 times a second, and not at all when no marker is in view. The AR.Drone itself reports its speed about 200 times a second. Adding up speed over time (dead reckoning) is smooth but slowly drifts. A Kalman filter is the standard way to combine the two: it predicts with the speed and corrects whenever a marker fix arrives, weighting each by how much it can be trusted.

In the paper’s filter, the state is the drone’s position (x, y, z), its heading (yaw) and the rates of change of all four. Heading is taken from the markers only. The drone reports speed in its own body frame, so it is first rotated into room coordinates using the heading:

Vx = vx·cos φ − vy·sin φ  // eq. 11 (φ = heading)
Vy = vx·sin φ + vy·cos φ  // eq. 12
Filter settingValue in the paper
Statex, y, z, yaw and their rates (8 numbers)
Measured by markersx, y, z and yaw
Prediction inputdrone speed, about 200 Hz
Correction inputmarker observation, about 10 Hz
Process noise Q0.1 on every diagonal entry
Measurement noise Rdiag(50, 50, 50, 0)

What this means: the filter trusts the speed-based prediction a lot between fixes, and pulls toward the marker when one is seen. The zero in R for yaw matches heading coming straight from the markers.

Animated. In the red gap no marker is in view: the estimate runs on speed alone and its uncertainty band widens, then snaps narrow when fixes return. (Illustrative numbers.)
Paper Fig. 11. A flight seen from above. Blue: position from markers alone, which jumps whenever the markers are lost. Green: the Kalman-filtered position, which the paper judges to be estimated reasonably well.
Paper Fig. 12. x, y and z over time while the drone was moved out of marker view part-way. Dots are marker readings; the line is the filtered estimate. Where the dots stop, the line continues from speed for a while, which is what lets the drone fly back to where it was after losing the markers.
Scene 5

Where next? Learning to choose

In short: in a simple grid simulator, the drone learns by trial and error where to stick markers, including what to do when a marker lands in the wrong place or falls off.

Deciding where to stick the next marker is not trivial: a marker may land off target, or fail to stick at all, and the plan has to adapt. The paper trains this choice with reinforcement learning: an agent tries actions, receives rewards, and gradually learns which action is worth most in each situation. It is trained offline in a simulator and the result is then used on the real drone.

The specific method is Q-learning. It keeps a score Q(s, a) for taking action a in situation s, and after every step nudges that score toward the reward just received plus the best score available from the next situation:

Q(s, a) ← Q(s, a) + α [ r + γ·max Q(s′, a′) − Q(s, a) ]  // eq. 1; α = learning rate, γ = discount
Animated. Each move costs a little battery; reaching the edge and widening the area pays off. Over three tries, the value of the good ending spreads back to earlier cells. (Toy numbers; the paper does not state α or γ.)

The simulator

Setting (Table 1)Value
Map size30 × 30 cells
Start point (where it can recharge)(15, 15)
Markers carried15
Battery capacity50
Maximum number of recharges5
Field of view5 cells
When sticking a marker (Table 2)Probability
Sticks as intended0.3
If not: lands shifted0.7
If not: is dropped0.3

What this means: the simulator was deliberately pessimistic. Most attempts go somewhat wrong, so the learned choices have to work even when markers do not end up where planned.

The situation the agent sees is its own position, its battery level and the map of markers. Its action is to move to a cell it can currently reach. Rewards: +3 for each unit of newly reachable area, +100 for each recharge, and −1 for each unit of battery used. It was trained for 100 steps.

Paper Fig. 1. The simulator’s map. s is the start and charging point, m are markers, and the grey region is where the drone can currently move.
Paper Fig. 2. Learning result. Left: a map state (green = markers, red = the drone). Right: the learned action values in that state. Values are higher toward the outside, so the agent tries to widen its area.
Paper Fig. 3. Reward over the 100 training steps. The paper notes that learning did not progress much: with so many possible situations, the same one almost never came up twice, so few scores got updated.

The authors’ conclusion is measured: training did not get far, but because the simulator included shifted markers, the learned values could still be used to pick marker positions that cope with such shifts. To make learning work better they suggest more steps, treating rotated or mirrored maps as the same situation to cut the number of states, and a better learning method.

Scene 6

Results on the real drone

In short: the pieces worked together on real hardware, and the drone stuck five markers in a row with no one at the controls.

Flying to a point

To stick a marker at the chosen spot, the drone must fly there and hold still. send_cmd uses a PD controller, one of the simplest feedback controllers. The speed command is a proportional part (how far away the target is) plus a derivative part (how fast that gap is changing):

u = Kp·Δx + Kd·(Δx_i − Δx_(i−1)) / Δt,  Kp = 0.15, Kd = 0.4  // eqs. 13–14; same for y and z
Animated. The command shrinks as the gap closes, and the derivative term brakes the approach. Gains are the paper’s; the drone’s response here is illustrative.
Paper Fig. 13. Position control toward the target (0, 0, −100) in the system’s coordinates (the paper does not state the unit). After take-off, control starts at the first annotation (制御開始) and ends at the second (制御終了) before landing. The paper judges the accuracy sufficient for a marker-sticking mission.

The full mission

ExperimentWhat the paper reports
Marker recognition (Sec. 4.1)Markers on the ceiling were recognized one after another; their positions and angles were registered.
Kalman filter (Sec. 4.2)Marker-only position jumps when markers are lost; the filtered position keeps a reasonable estimate from speed for a while.
Position control (Sec. 4.3)Reached the target point with accuracy judged enough for sticking markers.
Mission (Sec. 4.4)5 markers stuck in a row, automatically.

What this means: every link in the chain (see a marker, know where you are, fly to a spot, stick a new marker, register it) worked on real hardware, well enough to repeat five times without help.

Real run. Later in the same video: after moving on, the drone climbs back to the ceiling and hovers beside the markers already there to stick another one. Muted loop, cropped to hide the video’s labels. Footage: Takahiro Miki (author’s video), The University of Tokyo.
Paper Fig. 15. The ceiling after the mission. Markers that sit in clumps are ones the drone went to stick at the same place. The paper takes this as a sign that the positioning for sticking works, even though the choice of spot repeated.
Wrap-up

Limits and next steps

In short: a working proof of concept, with two clear open problems the authors name themselves.

  • The learner barely learned. There were too many possible situations, so scores rarely got updated. The authors suggest cutting the number of states (for example, treating rotated or mirrored maps as the same) and trying better learning methods.
  • Markers landed off target. The authors think the offset grew the shorter the drone could keep a marker in view while sticking. Their proposed fixes: better control during the sticking move, and learning to stick markers where they stay visible longer.
  • Scale of the demonstration. The mission covered one ceiling with five new markers; the paper does not report repeated trials or a size for the expanded area.

The idea itself, a robot that lays down its own landmarks as it explores instead of relying on a prepared environment, is what makes this short paper worth a look.