Guardrails
The safety limits built into an AI tool to keep it from producing harmful, dangerous, or inappropriate content, and to keep it acting within sensible bounds. Like the barriers on a mountain road, they are there to stop things going badly wrong. These built-in limits are called guardrails.
Guardrails are the rules and safeguards that shape how an AI behaves. They are put in place by the people who build and run the tool, and their job is to keep it helpful while stopping it from doing harm. That covers refusing dangerous requests, declining to produce hateful or explicit material, and, in tools that can take actions, holding back from anything risky without a check.
They take several forms working together. Some are trained into the model itself during its development. Some are written into a system prompt, the hidden set of instructions a tool follows behind the scenes. Others are separate filters that watch what goes in and what comes out. You rarely see any of this directly; you notice it mainly when an AI politely declines something.
Guardrails are a genuine good, and they are also imperfect. Set too loosely, they let harmful content slip through. Set too tightly, they frustrate people with needless refusals of reasonable requests. Getting the balance right is an ongoing piece of work, and different tools draw the line in different places.
You will also hear about people deliberately trying to get around these limits, which is called a jailbreak. For everyday use, it helps to see guardrails not as the AI being awkward, but as the reason you can hand these tools to a beginner, or a child, with a reasonable degree of trust.
Related terms
Used in
- AI safety and ethics · Guide