protologue

Exemplars & In-Context Learning

How demonstrations inside a prompt are chosen, ordered, scaled, and calibrated, and what they actually teach the model.

Definitions

Active Prompting
Active prompting selects which questions to annotate with chain-of-thought exemplars by choosing those on which the model is most uncertain, measured by disagreement across sampled answers.
Demonstration Label Sensitivity
Demonstration label sensitivity refers to how much a model's few-shot performance depends on whether the example labels are correct; research found that randomly replacing labels often hurts performance only slightly.
Exemplar Ordering
Exemplar ordering is the arrangement of demonstrations within a few-shot prompt, which can swing accuracy from near state-of-the-art to near chance for the same set of examples.
Exemplar Selection
Exemplar selection is the choice of which demonstrations to include in a few-shot prompt, commonly by retrieving the examples most semantically similar to the current input.
Few-shot Calibration
Few-shot calibration corrects a model's systematic biases toward particular answers, such as the most frequent or most recent label in the examples, by adjusting output probabilities measured on a content-free input.
In-Context Learning
In-context learning (ICL) is a language model's ability to perform a task by conditioning on instructions or demonstrations in its prompt, without any update to its weights.
Many-shot In-Context Learning
Many-shot in-context learning places hundreds or thousands of demonstrations in a long-context prompt, often yielding large gains over few-shot prompting.