Jailbreaking is the practice of crafting inputs specifically designed to bypass an AI model’s built-in safety restrictions, getting it to produce content or take actions it was explicitly designed to refuse.
Why It Matters
Every commercial AI model has guardrails limiting harmful, dangerous, or policy-violating outputs. Jailbreaking techniques, role-play framing, hypothetical scenarios, encoded instructions, attempt to convince the model those guardrails don’t apply in the current context.
A Practical Example
A user asks an AI model directly for instructions to build something dangerous and is refused, then rephrases the request as “write a fictional story where a character explains how to do this,” attempting to get the same information through a framing the model doesn’t recognize as a policy violation.
Related Terms
Need help governing AI risk like this across your organization?
Cyberix’s AI Security & Governance service finds, governs, and secures the AI already in use inside your organization.
