protologue

Constitutional AI

Also called CAI, RLAIF, reinforcement learning from AI feedback.

Constitutional AI is a training method in which a model critiques and revises its own outputs according to a written set of principles, and AI-generated preference judgments replace most human labels for harmlessness.

Description

Bai et al. used a supervised self-critique phase followed by reinforcement learning from AI feedback. The approach made the principles governing model behavior explicit and editable.

Sources

  1. Bai et al. (2022). Constitutional AI: Harmlessness from AI Feedback.

Cite this entry

Protologue. (2026). Constitutional AI. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0095). https://protologue.com/t/constitutional-ai/

BibTeX
@misc{protologue_constitutional_ai,
  title = {Constitutional AI},
  author = {{Protologue}},
  year = {2026},
  howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
  note = {Entry PTL-0095},
  url = {https://protologue.com/t/constitutional-ai/}
}

Markdown JSON