Constitutional AI
Also called CAI, RLAIF, reinforcement learning from AI feedback.
Constitutional AI is a training method in which a model critiques and revises its own outputs according to a written set of principles, and AI-generated preference judgments replace most human labels for harmlessness.
Description
Bai et al. used a supervised self-critique phase followed by reinforcement learning from AI feedback. The approach made the principles governing model behavior explicit and editable.
Sources
- Bai et al. (2022). Constitutional AI: Harmlessness from AI Feedback.
Cite this entry
Protologue. (2026). Constitutional AI. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0095). https://protologue.com/t/constitutional-ai/
BibTeX
@misc{protologue_constitutional_ai,
title = {Constitutional AI},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0095},
url = {https://protologue.com/t/constitutional-ai/}
}