Context Window Explained: Why AI Chatbots Forget What You Said Earlier

What a context window actually is, why it limits how much an AI model can remember in a conversation, and how token limits differ from a model's overall knowledge.
Anyone who has had a long conversation with an AI chatbot has probably noticed it eventually seems to “forget” something mentioned earlier. This isn’t the AI losing interest or malfunctioning; it’s running into the boundaries of its context window, a hard technical limit on how much text a model can actually consider at one time.
What a Context Window Actually Is
A context window is the maximum amount of text, measured in tokens rather than words, that a language model can process at once when generating a response. This includes everything currently relevant to the conversation: the system instructions, the entire visible conversation history, any documents or files provided, and the model’s own response as it’s being generated. Once a conversation’s total content exceeds the model’s context window size, the oldest portions typically have to be dropped, summarized, or otherwise excluded to make room, which is why a chatbot can seem to lose track of details mentioned much earlier in a long conversation.
Tokens, Not Words, Are the Real Unit of Measurement
Context window sizes are measured in tokens, not words or characters, where a token typically represents a common word, part of a longer word, or a punctuation mark. As a rough approximation, a token corresponds to about three-quarters of a word in English, meaning a context window advertised at 128,000 tokens can hold roughly 90,000 to 100,000 words, though this ratio varies somewhat by language and content type, with code and non-English text sometimes using more tokens per equivalent amount of content.
Why Context Window Size Doesn't Equal General Knowledge
A common point of confusion is treating context window size as the same thing as how much a model “knows.” A model’s general knowledge comes from its training data, encoded into its parameters during training and effectively fixed once training is complete. The context window is a completely separate, temporary working memory used only for the current conversation or task, more like the amount of paper someone can spread out on a desk at once rather than everything stored in their long-term memory. A model with a huge context window can process an entire lengthy document in one request, but that doesn’t expand what the model permanently “knows” beyond the conversation in which that document was provided.
Why Larger Context Windows Have Become a Major Competitive Focus
Expanding context window size has been one of the most actively competitive areas of improvement among AI model developers, since a larger context window unlocks genuinely new capabilities: analyzing entire lengthy documents or codebases in a single request, maintaining coherent, detailed context across very long conversations, and processing multiple documents together for comparison or synthesis. Larger context windows also come with real trade-offs, since processing more tokens generally requires more computing power and can increase both response time and cost, and some research has found that models don’t always use information from the middle of a very long context as reliably as information near the beginning or end.
Practical Tips for Working Within a Context Window
For tasks involving long documents or extended conversations, periodically summarizing key points, referencing specific important details explicitly rather than assuming the model still has them in mind, and starting a fresh conversation for a genuinely new, unrelated task all help work within a context window’s practical limits. Checking a specific AI tool’s stated context window size before relying on it for a very long document or an extended back-and-forth conversation avoids the frustration of watching earlier context quietly drop out of consideration.
Bottom Line
A context window is the temporary working memory limit that determines how much text an AI model can actually consider at once, separate from the permanent general knowledge baked into the model during training. Understanding this distinction explains why long conversations sometimes lose track of earlier details, and why context window size has become one of the most closely watched specifications when comparing different AI models.
Sources
- Academic research papers on transformer context length and attention mechanisms
- OpenAI, Anthropic and Google technical documentation on model context window specifications
- Independent research on long-context model performance and reliability
- Industry technical explainers on tokenization and context window measurement