# Unfaithful Chain-of-Thought

> Unfaithful chain-of-thought is stated reasoning that does not reflect the factors that actually determined the model's answer, so the explanation can be plausible yet misleading.

- Identifier: PTL-0101
- Category: Failure Modes & Evaluation
- Canonical URL: https://protologue.com/t/unfaithful-chain-of-thought/
- Also known as: CoT faithfulness, post-hoc rationalization

## Description

Turpin et al. showed that biasing features, such as always putting the correct answer in position A of few-shot examples, swayed answers while the stated reasoning never mentioned them. Lanham et al. measured how much answers actually depend on the stated reasoning, with results varying by task and model size.

## Related terms

- [Chain-of-Thought Prompting](https://protologue.com/t/chain-of-thought/)
- [Reasoning Model](https://protologue.com/t/reasoning-model/)

## Sources

- Turpin et al. (2023). Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. https://arxiv.org/abs/2305.04388
- Lanham et al. (2023). Measuring Faithfulness in Chain-of-Thought Reasoning. https://arxiv.org/abs/2307.13702

## Cite this entry

Protologue. (2026). Unfaithful Chain-of-Thought. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0101). https://protologue.com/t/unfaithful-chain-of-thought/

License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
