Can a chatbot keep a secret? Researchers turned that question into a public game: 163 teams guarded a password hidden in a chatbot’s instructions and tried to steal each other’s — and every guard was eventually broken.
Why this matters
AI assistants now answer customers, search the web and read email for us. The company gives each one hidden instructions (“only answer billing questions”) and sometimes private data. If a stranger can talk the assistant out of those instructions — say, with a line hidden in an email it reads — it can misbehave or leak that data. So: does telling the model “keep this secret” actually work?
What makes it hard
The model reads its instructions and a stranger’s words as one stream of text; there is no wall between the owner’s orders and what someone typed. It’s like a receptionist told “never give out the manager’s number” who then gets a caller with a very convincing story. And the guard can’t just refuse everything: an assistant that always says “I can’t help” is safe but useless.
What people did before
Researchers showed in 2022 that such attacks work, and online games collected them: HackAPrompt (organizers wrote the defenses, one message per try), Tensor Trust (players wrote defense instructions and attacked each other) and Gandalf (a secret word; most data kept private). But the defenses were only instructions and simple rules, and each attack was a single message.
What this paper does
A capture-the-flag contest at the IEEE SaTML 2024 security conference. The “flag” is a random 6-character password in the chatbot’s hidden instructions. Teams first built the strongest guard they could — extra instructions plus two automatic checks on every reply — that still kept the chatbot useful. Then they attacked everyone else’s guards, in full conversations, scoring for few attempts. Every chat was recorded.
What they showed
44 guards made it into the contest, and attackers ran 137,063 conversations against them. Every guard was broken at least once. Winners first worked out how a guard worked (spotting fake “decoy” passwords, for example), then got around it — like having the password spelled out in pieces the checks didn’t recognize. Persistence paid: 15 % of successful chats used four or more messages, against about 6 % of failed ones.
Why it’s a step forward
It is the first large public dataset where defenses include real output checks and attacks run over many turns — a ready-made test for future defenses — plus the open-source contest platform. The lesson for builders: instructions plus filters aren’t enough; test against attackers who adapt. The honest limit: it’s about one short secret and two 2024-era models.
- system prompt
- hidden instructions an app gives the chatbot before your chat starts
- prompt injection
- text written to make a chatbot ignore those instructions or leak them
- filter
- a small program or second AI that checks each reply before you see it
- decoy
- a fake secret planted to mislead attackers
- jailbreak
- getting a model to do what it was told or trained not to do
- utility
- how useful the chatbot stays for normal questions