Many-shot Jailbreaking
Many-shot jailbreaking fills a long context window with many fabricated dialogue examples in which an assistant complies with harmful requests, exploiting in-context learning to override the model's safety training.
Description
Anthropic researchers found the attack's effectiveness followed a power law in the number of shots, mirroring the scaling of benign in-context learning.
Sources
- Anthropic (2024). Many-shot jailbreaking.
Cite this entry
Protologue. (2026). Many-shot Jailbreaking. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0092). https://protologue.com/t/many-shot-jailbreaking/
BibTeX
@misc{protologue_many_shot_jailbreaking,
title = {Many-shot Jailbreaking},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0092},
url = {https://protologue.com/t/many-shot-jailbreaking/}
}