# Instruction Hierarchy

> The instruction hierarchy is a training approach that teaches a model to prioritize instructions by source, typically system over user over tool output, and to ignore lower-priority instructions that conflict with higher-priority ones.

- Identifier: PTL-0093
- Category: Security & Adversarial Prompting
- Canonical URL: https://protologue.com/t/instruction-hierarchy/
- Introduced: 2024

## Description

Wallace et al. trained models on synthetic conflicts and found large gains in robustness to prompt injection and system-prompt extraction with limited loss in helpfulness.

## Related terms

- [System Prompt](https://protologue.com/t/system-prompt/)
- [Prompt Injection](https://protologue.com/t/prompt-injection/)
- [Spotlighting](https://protologue.com/t/spotlighting/)

## Sources

- Wallace et al. (2024). The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions. https://arxiv.org/abs/2404.13208

## Cite this entry

Protologue. (2026). Instruction Hierarchy. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0093). https://protologue.com/t/instruction-hierarchy/

License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
