Tokenization Explained: Why AI Models Don’t Actually Read Words

How tokenization breaks text into the units AI language models actually process, why it affects cost and context limits, and why AI models sometimes miscount letters.
Large language models don’t process text as whole words or letters the way a human reader does. They work with tokens, a specific unit created through a process called tokenization, and understanding this step explains several otherwise puzzling AI quirks, including why AI models sometimes struggle with simple letter-counting tasks.
What Tokenization Actually Does
Tokenization is the process of breaking raw text down into smaller pieces, called tokens, before a language model can process it. A token might represent a whole common word, a piece of a longer or less common word, a single character, or a punctuation mark, depending on the specific tokenization method used and how frequently that particular piece of text appeared in the data used to build the tokenizer. Common words often become a single token, while rarer or more complex words, and especially words in less common languages, frequently get split into several smaller token pieces. This entire process happens before the language model itself ever sees the text, converting raw text into a numerical sequence the model’s underlying mathematics can actually process.
Why Tokenization Explains Some Odd AI Behavior
Because a language model processes tokens rather than individual letters, it doesn’t have direct, native access to a word’s exact letter-by-letter spelling the way a token-unaware system would. This is part of why AI models have historically sometimes struggled with tasks like counting the number of a specific letter within a word, since the word may be represented internally as one or two token chunks rather than a clear sequence of individual letters the model can straightforwardly count. This particular limitation has improved somewhat in more recent models through better training approaches, but it’s a good illustration of how a model’s internal representation of text can differ meaningfully from how a human perceives the same text.
Why Tokenization Affects Cost and Speed
Most commercial AI model providers charge for API access based on the number of tokens processed, both in the input prompt and the generated output, rather than a simple word or character count. This means the specific tokenizer a model uses can meaningfully affect real-world cost: a tokenizer that splits a given piece of text into more tokens results in higher cost and slower processing for that same content, compared with a more efficient tokenizer that represents the same text using fewer tokens. This is also directly connected to context window limits, since those limits are measured in tokens, meaning the specific tokenization method affects exactly how much actual text content fits within a given context window.
Why Different Languages Tokenize Differently
Because tokenizers are typically built based on patterns observed in large training datasets that skew heavily toward English and a handful of other widely represented languages, text in less-represented languages frequently gets broken into more, smaller tokens than equivalent English text would require. This has a real, practical consequence: processing the same amount of meaningful content in an underrepresented language can cost more and consume more of a model’s context window than equivalent English content, an issue that AI researchers and developers have increasingly worked to address through more balanced, multilingual tokenizer training.
Bottom Line
Tokenization converts raw text into the token units that language models actually process internally, a step that shapes API pricing, context window limits, and even certain unusual AI behaviors like letter-counting mistakes. Understanding that AI models work with these token chunks rather than directly with words or letters helps explain quirks that otherwise seem inconsistent with how capable these systems are in other respects.
Sources
- Academic research papers on subword tokenization methods (BPE, WordPiece, SentencePiece)
- OpenAI, Anthropic and Google technical documentation on tokenizer design
- Independent research on multilingual tokenization efficiency and fairness
- Industry technical explainers on token-based API pricing models