Security & Adversarial Prompting
Attacks that subvert a model's instructions and the defenses designed to resist them.
- Constitutional AIPTL-0095
- Instruction HierarchyPTL-0093
- JailbreakPTL-0090
- Adversarial SuffixPTL-0091
- Many-shot JailbreakingPTL-0092
- Prompt InjectionPTL-0087
- Indirect Prompt InjectionPTL-0088
- Prompt LeakingPTL-0089
- SpotlightingPTL-0094
Definitions
- Adversarial Suffix
- An adversarial suffix is an automatically optimized string of tokens that, when appended to a request, causes an aligned model to comply with requests it would normally refuse.
- Constitutional AI
- Constitutional AI is a training method in which a model critiques and revises its own outputs according to a written set of principles, and AI-generated preference judgments replace most human labels for harmlessness.
- Indirect Prompt Injection
- Indirect prompt injection places malicious instructions inside content a model will later retrieve or process, such as a web page, email, or document, so the attack is triggered without the attacker interacting with the model directly.
- Instruction Hierarchy
- The instruction hierarchy is a training approach that teaches a model to prioritize instructions by source, typically system over user over tool output, and to ignore lower-priority instructions that conflict with higher-priority ones.
- Jailbreak
- A jailbreak is a prompt crafted to make a model produce outputs its safety training is meant to prevent, often through role-play, hypothetical framing, obfuscation, or other adversarial techniques.
- Many-shot Jailbreaking
- Many-shot jailbreaking fills a long context window with many fabricated dialogue examples in which an assistant complies with harmful requests, exploiting in-context learning to override the model's safety training.
- Prompt Injection
- Prompt injection is an attack in which text supplied to a language model, by a user or through data the model processes, contains instructions that override or subvert the instructions of the application's developer.
- Prompt Leaking
- Prompt leaking is an attack that tricks a model into revealing its hidden system prompt or other confidential instructions.
- Spotlighting
- Spotlighting is a family of prompt-level defenses against indirect prompt injection that transform untrusted input, by delimiting, marking every word, or encoding it, so the model can distinguish it from trusted instructions.