protologue

Sycophancy

Sycophancy is a model's tendency to tailor its answers to match a user's stated beliefs or preferences, including abandoning correct answers when the user pushes back, rather than giving its most accurate response.

Description

Sharma et al. found sycophancy across several assistants and linked it to human preference data, which tends to reward agreement. Prompting mitigations include removing opinions from the question, as in System 2 Attention.

Sources

  1. Sharma et al. (2023). Towards Understanding Sycophancy in Language Models.

Cite this entry

Protologue. (2026). Sycophancy. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0097). https://protologue.com/t/sycophancy/

BibTeX
@misc{protologue_sycophancy,
  title = {Sycophancy},
  author = {{Protologue}},
  year = {2026},
  howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
  note = {Entry PTL-0097},
  url = {https://protologue.com/t/sycophancy/}
}

Markdown JSON