GenAI / LLM · Beginner Guide

How Do LLM Context Windows Work? Tokens, Limits and Why AI Forgets

What fills a model's working space, what happens when it overflows, and when a bigger window is the wrong fix.

Quick insight

Context window = the whiteboard, not the memory. The model only uses what is written on the board right now. When the board is full, something gets rubbed out, and the model is never told what.

An LLM context window is the maximum amount of text a large language model can work with in one request, measured in tokens. Your prompt, earlier messages, uploaded files, system instructions and the model's own answer all have to fit inside it. Anything outside the window does not exist for the model while it writes that answer.

GenAI / LLM

What is a context window in an LLM?

A context window is the model's active working space: everything it can read while producing one response. A large language model (LLM) is a model trained on huge amounts of text to predict the next piece of text; ChatGPT, Claude and Gemini are all built on LLMs.

Picture two people discussing a project in front of a whiteboard. Whatever is written on the board shapes what each person says next. As the discussion runs on, the board fills up, and older notes have to be rubbed out or squeezed into a summary. A context window works the same way.

Understanding this limit becomes especially useful when working with long conversations, large documents, AI agents, coding assistants, research workflows and enterprise AI applications.

Context is what gives a request its meaning. Compare these two prompts:

Example prompt · Prompt A
Suggest three marketing strategies.
Example prompt · Prompt B
My company sells industrial pumps.
Our main customers are manufacturing companies.
We want to increase sales in eastern India.
Suggest three marketing strategies.

Prompt B produces a far more useful answer only because the first three lines sit in the same context window as the request.

GenAI / LLM

What is a token, and how many tokens is a page of text?

A token is the unit of text a model actually reads: usually a whole word, part of a word, or a punctuation mark. A tokenizer, the component that splits text before the model sees it, might break “The customer wants faster delivery” into The · customer · wants · faster · delivery, and split a long word like “unbelievable” into two or three pieces. Context windows are measured in tokens, not words.

OpenAI's rule of thumb for English is that one token is roughly four characters, or about 0.75 words. That gives these working estimates:

TextApproximate wordsApproximate tokens
One email200270
One A4 page of prose500670
A 15-page report7,50010,000
A 200-page book100,000133,000

Three things push the count up: code, numbers and tables, and non-English text. Hindi, Bengali and other Indian languages often use noticeably more tokens per word than English with many tokenizers, so the same idea takes more of the window and, on paid APIs, costs more. Petrov and colleagues (NeurIPS 2023) measured this gap across languages and found the same sentence can need several times more tokens in some languages than in English. Paste a sample of your own text into a tokenizer tool to check rather than trusting the estimate.

GenAI / LLM

How does a context window work in practice?

The context window is a shared token budget: everything the application sends, plus the answer the model writes, draws from the same pool. Many models also cap the answer separately, at a smaller number.

A large language model generates its response from the information available inside that budget, which is why the content selected for the window directly affects the quality of the result.

Here is a simplified budget for a model with a 10,000-token window:

What the application sendsTokens
System instructions (rules the app gives the model)1,000
Conversation history (earlier messages, resent every turn)3,000
Uploaded document (text extracted from the file)4,000
Your current question500
Used8,500
Left for the answer1,500

Add a second 4,000-token document and the request no longer fits. The application, not the model, then decides what to do: drop older content, summarise it, or reject the request. Most chat apps do this silently.

Current context windows range from tens of thousands of tokens on smaller models to a million or more on some frontier models. These numbers change every few months, so check the vendor's documentation linked under Sources for the model you actually use.

GenAI / LLM

What takes up space in the context window?

Far more than your question. This is where most people underestimate how quickly a window fills.

ComponentWhat it isWhy it is easy to miss
System instructionsRules the app gives the model before you typeYou never see them
Conversation historyEarlier messages and repliesResent in full with every new turn
Uploaded filesText extracted from PDFs, spreadsheets, codeA “small” Excel file can run to tens of thousands of tokens
Retrieved passages (RAG)Snippets fetched from a knowledge base for this questionAdded behind the scenes
Tool and search resultsWeb pages, API responses, code outputOften long and unfiltered
The answer itselfEverything the model writes backLong answers use the same budget
Reasoning tokensInternal thinking on reasoning modelsHidden from you, but counted
Everything shares one budgetSystem 1,000Conversation 3,000Uploaded document 4,000Q 500Answer 1,50010,000-tokenlimit+ 2nd document 4,000Doesn't fit: the app drops, summarises or rejects — usually without telling you.
GenAI / LLM

What happens when a conversation exceeds the context window?

The application removes or compresses something, and the model then answers without it. The model does not know what is missing, and you are usually not told.

The simplest strategy is a sliding window: the app keeps only the most recent messages and silently drops the oldest.

A sliding window keeps only the latest messagesMessage 1Message 2“No scikit-learn, please”Message 3Message 4Message 5What the model sees (window of 3)Dropped. The model is not told.

Messages 1 and 2 have fallen outside the window. If message 1 said “Use only NumPy and Pandas, no scikit-learn”, the model will happily suggest scikit-learn later, because it can no longer see the rule. Real products use smarter strategies, such as summarising older turns or pinning important instructions, but something always gets left out.

Missing context also raises the risk of hallucination, which is when a model gives a confident answer that no source supports. If the sales report you uploaded has dropped out of the window and you ask for total revenue, the model may produce a plausible figure rather than say it no longer has the data. Hallucinations have other causes too, so check any number that matters against the original file.

GenAI / LLM

Why does a model miss information that is still inside the window?

Because fitting in the window is not the same as being used well. Two research findings explain this.

The first is the “Lost in the Middle” effect. Liu and colleagues (published in Transactions of the ACL, 2024) hid one relevant document among many irrelevant ones and moved it around the input. The models they tested answered best when the key information sat at the start or end, and worst when it sat in the middle. A 100-page report can fit comfortably and the model can still miss the figure on page 50.

The second is the gap between advertised and effective context length. NVIDIA's RULER benchmark (Hsieh et al., 2024) tested models that claimed windows of 32,000 tokens or more and found that only about half kept satisfactory performance at 32,000 tokens. The number on the spec sheet is the most the model will accept, not the length at which it stays reliable.

Both effects vary by model and have improved in newer releases, but neither has disappeared.

Rule of thumb

Put critical instructions at the top of the prompt and repeat the actual question at the very end. Never leave the one thing that matters buried in the middle of a long document.

GenAI / LLM

Is a bigger context window always better?

No. A bigger window lets you supply more, but it costs more, runs slower and adds risk.

Compute grows fast. Transformer models, the architecture behind today's LLMs, use self-attention: every token is compared with every other token to work out how they relate. That is how a model works out who “she” is in “Riya sent the proposal to Neha because she asked for it”. In standard attention, the number of comparisons grows with the square of the input length:

Input length (tokens)Token-pair comparisons
1,0001 million
2,0004 million
4,00016 million

Techniques such as FlashAttention and caching make this far cheaper in practice, but long inputs still mean higher bills and slower answers.

More text means more ways to be attacked. Long contexts often contain web pages, emails and documents you did not write. Prompt injection is when hidden text in that content tries to override your instructions. OWASP ranks it first among risks for LLM applications. The more untrusted text you pour into the window, the more chances it has to work.

Honest limits

When don't you need a bigger context window?

Most of the time. In our experience training teams on generative AI, people reach for a longer window when a better workflow would be faster, cheaper and more accurate.

If your situation is…A bigger window is…Do this instead
Analysing 50,000 rows of sales data✗ The wrong toolCompute totals and trends in Python, SQL or Excel; give the model the summary table
Answering questions from hundreds of policy documents✗ Expensive and less accurateUse retrieval (RAG) to send only the relevant passages
A long chat that has started drifting✗ A temporary patchAsk for a summary of decisions; start a fresh chat with it
Summarising a 200-page report✗ Workable but riskySummarise section by section with page references, then combine
Reviewing one contract or one code module end to end✓ Genuinely usefulUse the long window; put the key question at the end
Strong opinion

Start with less context, not more. A long window earns its cost only when the model must see everything at once to reason across it. When it needs only the relevant parts, retrieving or computing those parts wins on cost, speed and accuracy.

GenAI / LLM

How is a context window different from memory, training data and RAG?

The context window is what the model can see right now. Training data is knowledge built in beforehand, and memory and RAG are ways of getting information into the window.

ConceptWhat it isHow long it lastsExample
Context windowText available to the model in this requestOne requestYour prompt plus the history the app resends
Training knowledgePatterns learned when the model was builtFixed until retrainedKnowing Python syntax
Persistent memoryFacts an app saves between conversationsUntil deletedPrefers weekly summary reports
RAGPassages fetched from a knowledge base on demandFetched per questionThe relevant clause from an exam policy

Memory and RAG do not get around the window. Whatever they retrieve still has to be placed into it, and still uses tokens, before it can affect the answer.

Here is how memory actually works. An app saves a fact about you, say “Arjun prefers a weekly summary report”, in its own database. The model never holds it. When you next ask for a report, the app pastes that line into your prompt behind the scenes, and only then does it become context. RAG works the same way, with a search step choosing which passages to paste in.

For when to use retrieval and when to retrain a model instead, see RAG vs fine-tuning.

Worked example

A worked example: twelve monthly sales reports

A retail analyst uploads twelve monthly sales reports, each 15 pages long, to a model with a 128,000-token window, and asks: “Which region grew fastest this year, and in which month did growth stall?”

ItemTokens
12 reports × 15 pages × ~670 tokens~120,000
System instructions1,500
Question500
Input total~122,000
Left for the answer~6,000

On paper it fits. In practice three things go wrong. The June to September reports sit in the middle of the input, where models are least reliable. Every follow-up question adds history and pushes the earliest reports out. And there is little room left for a detailed answer.

The better approach is to extract the regional figures into a table with Pandas or Excel, which takes a few minutes, and send the model a summary of about 2,000 tokens with the question. Input drops by roughly 98%, the answer is cheaper, faster and checkable, and the model spends its effort on interpretation, which is what it is good at.

GenAI / LLM

How can you use a context window more effectively?

Send less, structure it clearly, and put what matters at the edges.

  • Measure one real prompt: paste a document you regularly use with AI into a tokenizer tool and note the count.
  • Keep a project brief: goals, constraints and decisions so far. Paste it at the start of each new chat instead of letting one chat run for hours.
  • Label long prompts with headings: OBJECTIVE, BACKGROUND, DATA, CONSTRAINTS, REQUIRED OUTPUT.
  • Put critical instructions first and repeat the question last.
  • Compute numbers in code or Excel; send the model results, not raw rows.
  • Break large tasks into stages: extract, categorise, find patterns, recommend, check against the source.
  • Treat text you did not write, including web pages, emails and PDFs, as untrusted.
FAQ

Frequently asked questions

What is a context window in simple terms?

It is the amount of text an AI model can look at in one go, counted in tokens. It includes your question, earlier messages, any files attached and the answer the model writes. Anything beyond that limit is invisible to the model for that request.

What happens when ChatGPT reaches its context limit?

The app shortens what it sends to the model, usually by dropping or summarising older messages, or it refuses an input that is too large. You are rarely told which. The visible symptom is the model forgetting earlier instructions or details.

Is one token the same as one word?

No. A token can be a whole word, part of a word or a punctuation mark. In English, 100 tokens is roughly 75 words; code, numbers and Indian languages usually need more tokens for the same content.

Does a larger context window make an LLM smarter?

No. It lets you give the model more material, but reasoning quality is a separate property. Research such as Lost in the Middle and RULER shows models use long inputs unevenly, missing details in the middle or weakening well before their advertised limit.

Does RAG remove the context-window limit?

No. RAG chooses which passages to send so you do not have to send everything, but the retrieved text still occupies the window. It makes a limited window go much further; it does not make it unlimited.

Is ChatGPT's memory the same as its context window?

No. Memory stores selected facts across conversations; the context window is what the model can see in the current request. A saved memory only affects an answer once the app inserts it into the context window.

Sources

Sources

Keep reading

Related articles

More in GenAI / LLM.

Find your knowledge gaps on generative AI. Start your PrepAI Diagnose: a short quiz on this topic, with feedback on what to study next.

Start your PrepAI Diagnose

Want to build applications that handle long documents properly, with retrieval, summarisation pipelines and evaluation? Our Advanced Generative AI Course covers it hands-on. WhatsApp or call +91 7676882222.