Failure Modes & Evaluation
Systematic ways prompted models go wrong, and the evaluations used to measure them.
- HallucinationPTL-0096
- Lost in the MiddlePTL-0098
- Needle in a HaystackPTL-0099
- Prompt SensitivityPTL-0100
- SycophancyPTL-0097
- Unfaithful Chain-of-ThoughtPTL-0101
Definitions
- Hallucination
- Hallucination is generated content that is fluent and plausible but unfaithful to the provided source or factually incorrect, such as fabricated citations, facts, or quotations.
- Lost in the Middle
- Lost in the middle is the finding that language models use information at the beginning or end of a long context much more reliably than information placed in the middle.
- Needle in a Haystack
- Needle in a haystack is a long-context evaluation that hides a specific fact at varying depths in a long distractor document and tests whether the model can retrieve it.
- Prompt Sensitivity
- Prompt sensitivity is the variation in a model's performance caused by superficial changes to a prompt, such as formatting, separators, spacing, or wording, that do not change the task's meaning.
- Sycophancy
- Sycophancy is a model's tendency to tailor its answers to match a user's stated beliefs or preferences, including abandoning correct answers when the user pushes back, rather than giving its most accurate response.
- Unfaithful Chain-of-Thought
- Unfaithful chain-of-thought is stated reasoning that does not reflect the factors that actually determined the model's answer, so the explanation can be plausible yet misleading.