# Process Reward Model

> A process reward model (PRM) scores each intermediate step of a model's reasoning, rather than only the final answer, and is used to select or train better reasoning chains.

- Identifier: PTL-0053
- Category: Self-Critique & Verification
- Canonical URL: https://protologue.com/t/process-reward-model/
- Also known as: PRM, step-level verifier, process supervision
- Introduced: 2023

## Description

Lightman et al. showed process supervision outperformed outcome supervision for selecting correct solutions to competition math problems, and released a large dataset of step-level human labels.

## Related terms

- [Best-of-N Sampling](https://protologue.com/t/best-of-n-sampling/)
- [Test-Time Compute Scaling](https://protologue.com/t/test-time-compute-scaling/)
- [LLM-as-a-Judge](https://protologue.com/t/llm-as-a-judge/)

## Sources

- Lightman et al. (2023). Let's Verify Step by Step. https://arxiv.org/abs/2305.20050

## Cite this entry

Protologue. (2026). Process Reward Model. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0053). https://protologue.com/t/process-reward-model/

License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
