protologue

Test-Time Compute Scaling

Also called inference-time scaling, test-time scaling.

Test-time compute scaling improves a model's answers by spending more computation at inference, through longer reasoning, more samples, search, or verification, rather than by training a larger model.

Description

Snell et al. found that allocating test-time compute adaptively per prompt could be more effective than scaling model parameters for some problems. Self-consistency, best-of-N, tree search, and reasoning models are all forms of test-time scaling.

Sources

  1. Snell et al. (2024). Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Cite this entry

Protologue. (2026). Test-Time Compute Scaling. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0084). https://protologue.com/t/test-time-compute-scaling/

BibTeX
@misc{protologue_test_time_compute_scaling,
  title = {Test-Time Compute Scaling},
  author = {{Protologue}},
  year = {2026},
  howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
  note = {Entry PTL-0084},
  url = {https://protologue.com/t/test-time-compute-scaling/}
}

Markdown JSON