What Is a Context Window in LLMs?
You're twenty messages into a working session with an AI chatbot. Back in message two, you gave it a rule: keep every code example under thirty lines. Around message twenty, it hands you a ninety-line block. You repeat the rule. It apologizes, complies — and does it again an hour later.
The model didn't get careless. What happened is the closest thing an LLM has to forgetting: its context window filled up, and your rule from message two got quietly dropped to make room. Understanding that mechanism changes how you use these tools, so this article explains what a context window is, what 128K actually means in pages and words, why long chats degrade, and how to read the numbers vendors advertise.
What a context window is
A context window is the maximum amount of text — measured in tokens — that a model can process in a single inference pass. IBM likens it to working memory: it's everything the model can "see" at once while generating its next reply, and anything outside it simply doesn't exist for the model.
I find a desk metaphor more useful, because desks explain what goes wrong. Picture the model's context as a desktop with fixed surface area. Everything it works with — your question, the conversation so far, any attached documents — has to lie on that desk at the same time. Surface area runs out? Something gets taken off the desk before anything new goes on. There's no drawer, no filing cabinet. The window isn't the model's memory of you; it's the desk space for this one request.
One detail that definition pages mention once and hurry past: the window covers input and output combined. Your prompt, the model's in-progress answer, everything — all from one pool.
How big is 4K or 128K, really
Context windows are counted in tokens, and a token is the smallest unit of text a model reads — usually a word fragment, sometimes a whole word. There's no fixed exchange rate, but rough numbers hold: one English word ≈ 1.3 tokens, one Chinese character ≈ 0.6 tokens, per the tokenization walkthroughs from IBM and Chinese practitioner writeups. IBM's researchers also found tokenization isn't fair across languages — the same sentence in Telugu consumed seven times more tokens than English, so the "same" window holds less text in some languages.
Translate the jargon into pages and the marketing numbers get honest:
| Window | ≈ English words | ≈ What it holds |
|---|---|---|
| 2K | 1,500 | A long email thread (McKinsey notes this was GPT-3's limit) |
| 8K | 6,000 | A short paper or a few files of code |
| 32K | 24,000 | A whitepaper with appendix |
| 128K | 96,000 | A 200-page book |
| 1M–2M | 750K–1.5M | 1,500–3,000 pages — Gemini territory |
That first row explains a lot of history: GPT-3 shipped with 2,048 tokens, about 1,500 words — not enough for an insurance plan or a vendor contract. Everything below 8K now feels prehistoric, but as recently as 2023, entire workarounds existed because of that row.
The window holds more than your input
Ask for "128K context" and you might imagine 96,000 words of your documents. Here's what actually occupies the desk:
| Occupant | What it is |
|---|---|
| Your current input | The prompt, pasted documents |
| Conversation history | Every prior turn, resent each time |
| System prompt | Hidden instructions from the provider shaping behavior |
| Retrieved material | Documents a RAG pipeline fetched to answer you |
| Formatting | Special characters, line breaks, markup |
| The model's output | The answer being generated, drawn from the same pool |
IBM's explainer is one of the few that spells this out: system prompts and RAG retrievals live in the context window too. That's why an advertised window never feels as roomy as it sounds — by the time the visible parts arrive, the invisible parts have already claimed seats.
Why long chats "forget"
Now the message-twenty mystery. Chat APIs are stateless: the server doesn't remember your conversation. Every turn, the client resends the entire history — your messages, the model's replies, all of it — plus the new message. Practitioner guides to the DeepSeek API walk through this explicitly, with code: round two's request literally contains round one's question and answer.
So each turn, the desk gets fuller. When the pile finally exceeds the window, the provider's engineering kicks in: context truncation. The system keeps the most recent N tokens and silently drops the oldest content. You don't see an error — you see a model that can still chat fluently but has lost track of what it promised you fifty messages ago. That's not the model misremembering; it's the material no longer being on the desk at all.
A Chinese-language experiment on Tencent Cloud's developer community demonstrates the failure mode crisply: feed a 4,096-token model a 5,000-token product manual, and the first ~900 tokens get cut — the section on core features vanishes while the warranty section (at the end) survives. Ask about core features, you get garbage or a guess; ask about the warranty, a perfect answer. Same document, opposite outcomes, decided purely by position.
The practical fixes are unglamorous: start a fresh session once the task changes, restate load-bearing constraints near your new question instead of trusting turn two, and keep genuinely important reference material in the latest message. I've made "paste the requirements again" a habit before every major deliverable in a long session — cheap insurance against a rule that fell off the desk.
Advertised vs. effective: the number isn't the number
Here's where vendor specs and user experience diverge. A model may accept 128K tokens and still use the middle of that context poorly. The landmark study is "Lost in the Middle" (Liu et al., 2023): models retrieve information from the beginnings and ends of long inputs well, but performance dips when the answer sits mid-context. Wikipedia's editors track the corollary under the name effective context length — the point where performance actually starts deteriorating, which research finds is often shorter than the advertised window.
Vendors are aware and things are improving — Google DeepMind reported in 2024 that newer models handle long-range coherence better. But my rule of thumb when a tool "knows" something and still answers wrong: before blaming the model's intelligence, check whether the fact was buried mid-document. It's also why answers grounded in retrieved evidence beat raw memory — retrieval narrows what sits on the desk, which is one reason it also softens hallucination.
Why windows don't just grow forever
If bigger is better, why not ship a 100M-token window tomorrow? Math. The self-attention mechanism at the heart of transformers compares every token with every other token, so compute scales with the square of the length: double the tokens, quadruple the work. Community engineers estimate the attention matrix for a 128K sequence needs on the order of 65GB of memory in FP16 — versus 262MB at 8K. Longer contexts are also slower (each generated word re-references everything before it), more expensive on token-based pricing, and, per IBM citing Anthropic's research, a longer attack surface for jailbreaks.
Solutions exist — sparse attention, sliding windows, rotary position embeddings — and they're why windows grew from 4K to 1M+ in three years. But the constraint is real, which is why context length remains a headline spec rather than a solved problem.
The 2026 landscape and how to choose
Current advertised windows from public model documentation:
| Model family | Advertised window |
|---|---|
| OpenAI GPT-4o / o-series | 128K |
| Anthropic Claude | 200K standard (500K on enterprise plans) |
| Google Gemini | 1M–2M |
| Meta Llama 3.1+ | 128K |
| DeepSeek | 64K |
| Mistral Large | 128K |
Read that column as ceiling, not capacity — remember effective length runs shorter. For choosing: everyday chat and single-file questions fit comfortably in 8–32K. Whole contracts, papers, or a repo subtree want 128K+. Book-scale or cross-document analysis is 1M+ territory. And if you're picking a model mostly for its giant window, pause: for large private knowledge bases, retrieval (RAG) with a modest window usually beats stuffing everything into context — cheaper, auditable, and immune to the middle-of-the-document blind spot. The window decides how much the model can see at once; retrieval decides what's worth seeing.
Quick answers
What is a context window in simple terms? The maximum text an LLM can consider in one request — its working desk for that single exchange, covering input and output together.
Does the model's output count against the window? Yes. Input and output draw from the same pool; a model with 64K context and 8K max output leaves you 56K of input room.
Why does the AI forget things I told it earlier? Each turn resends the whole history, and once it exceeds the window, the oldest content is silently truncated. Restate critical constraints in your newest message.
Is a bigger context window always better? No. Beyond cost and speed, effective length often falls short of the advertised number — benchmarked performance degrades on mid-context information. For big document sets, RAG with a reasonable window usually wins.
Which model has the largest context window? As of 2026, Google's Gemini line advertises up to 1M–2M tokens (~1,500–3,000 pages). Kimi and a few others compete in the million-token range; check current docs, since these numbers move quarterly.
So the next time a launch headline brags about token counts, you'll know what a bigger context window actually buys you: more desk space, advertised at full width, in a world where mid-desk items still get overlooked. Judge models by what fits your documents — and when the pile outgrows any desk, that's the moment retrieval earns its keep.