Jailbreak
A deliberate attempt to trick an AI into ignoring its own safety rules and doing something it is meant to refuse. People try clever wording, role-play scenarios, or step-by-step manipulation to talk the model past its limits. Getting around an AI's guardrails in this way is called a jailbreak.
A jailbreak is an effort to make an AI break its own rules. Every well-built AI tool has guardrails that stop it from producing harmful or inappropriate content. A jailbreak tries to slip past those limits through the words of a prompt, persuading the model to do what it would normally decline.
The methods can be surprisingly creative. Someone might wrap a forbidden request in a story, claim it is “only hypothetical,” pretend to be a researcher who needs the information, or build up to the request in small steps so no single message looks alarming. The common thread is dressing up a blocked request until it slips through.
The term is borrowed from phones, where “jailbreaking” meant removing the manufacturer’s restrictions. It is closely related to prompt injection, though the two differ in aim. A jailbreak targets the model’s built-in rules directly through what you type; prompt injection smuggles instructions in through outside content the AI happens to read.
For an everyday user, this is mostly useful to understand rather than to attempt. Jailbreaking sits against the terms of use of most AI tools and against the whole point of having safety limits. Knowing the word helps you follow news stories about AI safety, where jailbreaks come up often as researchers test how sturdy these protections really are.