Token
Also called subword, BPE token.
A token is the basic unit of text a language model reads and writes, typically a word, word fragment, or character sequence produced by a subword tokenizer such as byte-pair encoding.
Description
Context limits, pricing, and generation speed are all measured in tokens. Because tokenization splits text unevenly, character-level tasks such as counting letters or reversing strings are harder for models than they appear, and the same content can cost different numbers of tokens in different languages.
Sources
- Sennrich et al. (2015). Neural Machine Translation of Rare Words with Subword Units.
Cite this entry
Protologue. (2026). Token. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0005). https://protologue.com/t/token/
BibTeX
@misc{protologue_token,
title = {Token},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0005},
url = {https://protologue.com/t/token/}
}