protologue

Security & Adversarial Prompting

Attacks that subvert a model's instructions and the defenses designed to resist them.

Definitions

Adversarial Suffix
An adversarial suffix is an automatically optimized string of tokens that, when appended to a request, causes an aligned model to comply with requests it would normally refuse.
Constitutional AI
Constitutional AI is a training method in which a model critiques and revises its own outputs according to a written set of principles, and AI-generated preference judgments replace most human labels for harmlessness.
Indirect Prompt Injection
Indirect prompt injection places malicious instructions inside content a model will later retrieve or process, such as a web page, email, or document, so the attack is triggered without the attacker interacting with the model directly.
Instruction Hierarchy
The instruction hierarchy is a training approach that teaches a model to prioritize instructions by source, typically system over user over tool output, and to ignore lower-priority instructions that conflict with higher-priority ones.
Jailbreak
A jailbreak is a prompt crafted to make a model produce outputs its safety training is meant to prevent, often through role-play, hypothetical framing, obfuscation, or other adversarial techniques.
Many-shot Jailbreaking
Many-shot jailbreaking fills a long context window with many fabricated dialogue examples in which an assistant complies with harmful requests, exploiting in-context learning to override the model's safety training.
Prompt Injection
Prompt injection is an attack in which text supplied to a language model, by a user or through data the model processes, contains instructions that override or subvert the instructions of the application's developer.
Prompt Leaking
Prompt leaking is an attack that tricks a model into revealing its hidden system prompt or other confidential instructions.
Spotlighting
Spotlighting is a family of prompt-level defenses against indirect prompt injection that transform untrusted input, by delimiting, marking every word, or encoding it, so the model can distinguish it from trusted instructions.