An LLM context window is the maximum amount of text a large language model can work with in one request, measured in tokens. Your prompt, earlier messages, uploaded files, system instructions and the model's own answer all have to fit inside it. Anything outside the window does not exist for the model while it writes that answer.
What is a context window in an LLM?
A context window is the model's active working space: everything it can read while producing one response. A large language model (LLM) is a model trained on huge amounts of text to predict the next piece of text; ChatGPT, Claude and Gemini are all built on LLMs.
Picture two people discussing a project in front of a whiteboard. Whatever is written on the board shapes what each person says next. As the discussion runs on, the board fills up, and older notes have to be rubbed out or squeezed into a summary. A context window works the same way.
Understanding this limit becomes especially useful when working with long conversations, large documents, AI agents, coding assistants, research workflows and enterprise AI applications.
Context is what gives a request its meaning. Compare these two prompts:
Our main customers are manufacturing companies.
We want to increase sales in eastern India.
Suggest three marketing strategies.
Prompt B produces a far more useful answer only because the first three lines sit in the same context window as the request.
What is a token, and how many tokens is a page of text?
A token is the unit of text a model actually reads: usually a whole word, part of a word, or a punctuation mark. A tokenizer, the component that splits text before the model sees it, might break “The customer wants faster delivery” into The · customer · wants · faster · delivery, and split a long word like “unbelievable” into two or three pieces. Context windows are measured in tokens, not words.
OpenAI's rule of thumb for English is that one token is roughly four characters, or about 0.75 words. That gives these working estimates:
| Text | Approximate words | Approximate tokens |
|---|---|---|
| One email | 200 | 270 |
| One A4 page of prose | 500 | 670 |
| A 15-page report | 7,500 | 10,000 |
| A 200-page book | 100,000 | 133,000 |
Three things push the count up: code, numbers and tables, and non-English text. Hindi, Bengali and other Indian languages often use noticeably more tokens per word than English with many tokenizers, so the same idea takes more of the window and, on paid APIs, costs more. Petrov and colleagues (NeurIPS 2023) measured this gap across languages and found the same sentence can need several times more tokens in some languages than in English. Paste a sample of your own text into a tokenizer tool to check rather than trusting the estimate.
How does a context window work in practice?
The context window is a shared token budget: everything the application sends, plus the answer the model writes, draws from the same pool. Many models also cap the answer separately, at a smaller number.
A large language model generates its response from the information available inside that budget, which is why the content selected for the window directly affects the quality of the result.
Here is a simplified budget for a model with a 10,000-token window:
| What the application sends | Tokens |
|---|---|
| System instructions (rules the app gives the model) | 1,000 |
| Conversation history (earlier messages, resent every turn) | 3,000 |
| Uploaded document (text extracted from the file) | 4,000 |
| Your current question | 500 |
| Used | 8,500 |
| Left for the answer | 1,500 |
Add a second 4,000-token document and the request no longer fits. The application, not the model, then decides what to do: drop older content, summarise it, or reject the request. Most chat apps do this silently.
Current context windows range from tens of thousands of tokens on smaller models to a million or more on some frontier models. These numbers change every few months, so check the vendor's documentation linked under Sources for the model you actually use.
What takes up space in the context window?
Far more than your question. This is where most people underestimate how quickly a window fills.
| Component | What it is | Why it is easy to miss |
|---|---|---|
| System instructions | Rules the app gives the model before you type | You never see them |
| Conversation history | Earlier messages and replies | Resent in full with every new turn |
| Uploaded files | Text extracted from PDFs, spreadsheets, code | A “small” Excel file can run to tens of thousands of tokens |
| Retrieved passages (RAG) | Snippets fetched from a knowledge base for this question | Added behind the scenes |
| Tool and search results | Web pages, API responses, code output | Often long and unfiltered |
| The answer itself | Everything the model writes back | Long answers use the same budget |
| Reasoning tokens | Internal thinking on reasoning models | Hidden from you, but counted |
What happens when a conversation exceeds the context window?
The application removes or compresses something, and the model then answers without it. The model does not know what is missing, and you are usually not told.
The simplest strategy is a sliding window: the app keeps only the most recent messages and silently drops the oldest.
Messages 1 and 2 have fallen outside the window. If message 1 said “Use only NumPy and Pandas, no scikit-learn”, the model will happily suggest scikit-learn later, because it can no longer see the rule. Real products use smarter strategies, such as summarising older turns or pinning important instructions, but something always gets left out.
Missing context also raises the risk of hallucination, which is when a model gives a confident answer that no source supports. If the sales report you uploaded has dropped out of the window and you ask for total revenue, the model may produce a plausible figure rather than say it no longer has the data. Hallucinations have other causes too, so check any number that matters against the original file.
Why does a model miss information that is still inside the window?
Because fitting in the window is not the same as being used well. Two research findings explain this.
The first is the “Lost in the Middle” effect. Liu and colleagues (published in Transactions of the ACL, 2024) hid one relevant document among many irrelevant ones and moved it around the input. The models they tested answered best when the key information sat at the start or end, and worst when it sat in the middle. A 100-page report can fit comfortably and the model can still miss the figure on page 50.
The second is the gap between advertised and effective context length. NVIDIA's RULER benchmark (Hsieh et al., 2024) tested models that claimed windows of 32,000 tokens or more and found that only about half kept satisfactory performance at 32,000 tokens. The number on the spec sheet is the most the model will accept, not the length at which it stays reliable.
Both effects vary by model and have improved in newer releases, but neither has disappeared.
Put critical instructions at the top of the prompt and repeat the actual question at the very end. Never leave the one thing that matters buried in the middle of a long document.
Is a bigger context window always better?
No. A bigger window lets you supply more, but it costs more, runs slower and adds risk.
Compute grows fast. Transformer models, the architecture behind today's LLMs, use self-attention: every token is compared with every other token to work out how they relate. That is how a model works out who “she” is in “Riya sent the proposal to Neha because she asked for it”. In standard attention, the number of comparisons grows with the square of the input length:
| Input length (tokens) | Token-pair comparisons |
|---|---|
| 1,000 | 1 million |
| 2,000 | 4 million |
| 4,000 | 16 million |
Techniques such as FlashAttention and caching make this far cheaper in practice, but long inputs still mean higher bills and slower answers.
More text means more ways to be attacked. Long contexts often contain web pages, emails and documents you did not write. Prompt injection is when hidden text in that content tries to override your instructions. OWASP ranks it first among risks for LLM applications. The more untrusted text you pour into the window, the more chances it has to work.
When don't you need a bigger context window?
Most of the time. In our experience training teams on generative AI, people reach for a longer window when a better workflow would be faster, cheaper and more accurate.
| If your situation is… | A bigger window is… | Do this instead |
|---|---|---|
| Analysing 50,000 rows of sales data | ✗ The wrong tool | Compute totals and trends in Python, SQL or Excel; give the model the summary table |
| Answering questions from hundreds of policy documents | ✗ Expensive and less accurate | Use retrieval (RAG) to send only the relevant passages |
| A long chat that has started drifting | ✗ A temporary patch | Ask for a summary of decisions; start a fresh chat with it |
| Summarising a 200-page report | ✗ Workable but risky | Summarise section by section with page references, then combine |
| Reviewing one contract or one code module end to end | ✓ Genuinely useful | Use the long window; put the key question at the end |
Start with less context, not more. A long window earns its cost only when the model must see everything at once to reason across it. When it needs only the relevant parts, retrieving or computing those parts wins on cost, speed and accuracy.
How is a context window different from memory, training data and RAG?
The context window is what the model can see right now. Training data is knowledge built in beforehand, and memory and RAG are ways of getting information into the window.
| Concept | What it is | How long it lasts | Example |
|---|---|---|---|
| Context window | Text available to the model in this request | One request | Your prompt plus the history the app resends |
| Training knowledge | Patterns learned when the model was built | Fixed until retrained | Knowing Python syntax |
| Persistent memory | Facts an app saves between conversations | Until deleted | Prefers weekly summary reports |
| RAG | Passages fetched from a knowledge base on demand | Fetched per question | The relevant clause from an exam policy |
Memory and RAG do not get around the window. Whatever they retrieve still has to be placed into it, and still uses tokens, before it can affect the answer.
Here is how memory actually works. An app saves a fact about you, say “Arjun prefers a weekly summary report”, in its own database. The model never holds it. When you next ask for a report, the app pastes that line into your prompt behind the scenes, and only then does it become context. RAG works the same way, with a search step choosing which passages to paste in.
For when to use retrieval and when to retrain a model instead, see RAG vs fine-tuning.
A worked example: twelve monthly sales reports
A retail analyst uploads twelve monthly sales reports, each 15 pages long, to a model with a 128,000-token window, and asks: “Which region grew fastest this year, and in which month did growth stall?”
| Item | Tokens |
|---|---|
| 12 reports × 15 pages × ~670 tokens | ~120,000 |
| System instructions | 1,500 |
| Question | 500 |
| Input total | ~122,000 |
| Left for the answer | ~6,000 |
On paper it fits. In practice three things go wrong. The June to September reports sit in the middle of the input, where models are least reliable. Every follow-up question adds history and pushes the earliest reports out. And there is little room left for a detailed answer.
The better approach is to extract the regional figures into a table with Pandas or Excel, which takes a few minutes, and send the model a summary of about 2,000 tokens with the question. Input drops by roughly 98%, the answer is cheaper, faster and checkable, and the model spends its effort on interpretation, which is what it is good at.
How can you use a context window more effectively?
Send less, structure it clearly, and put what matters at the edges.
- Measure one real prompt: paste a document you regularly use with AI into a tokenizer tool and note the count.
- Keep a project brief: goals, constraints and decisions so far. Paste it at the start of each new chat instead of letting one chat run for hours.
- Label long prompts with headings: OBJECTIVE, BACKGROUND, DATA, CONSTRAINTS, REQUIRED OUTPUT.
- Put critical instructions first and repeat the question last.
- Compute numbers in code or Excel; send the model results, not raw rows.
- Break large tasks into stages: extract, categorise, find patterns, recommend, check against the source.
- Treat text you did not write, including web pages, emails and PDFs, as untrusted.
Frequently asked questions
What is a context window in simple terms?
It is the amount of text an AI model can look at in one go, counted in tokens. It includes your question, earlier messages, any files attached and the answer the model writes. Anything beyond that limit is invisible to the model for that request.
What happens when ChatGPT reaches its context limit?
The app shortens what it sends to the model, usually by dropping or summarising older messages, or it refuses an input that is too large. You are rarely told which. The visible symptom is the model forgetting earlier instructions or details.
Is one token the same as one word?
No. A token can be a whole word, part of a word or a punctuation mark. In English, 100 tokens is roughly 75 words; code, numbers and Indian languages usually need more tokens for the same content.
Does a larger context window make an LLM smarter?
No. It lets you give the model more material, but reasoning quality is a separate property. Research such as Lost in the Middle and RULER shows models use long inputs unevenly, missing details in the middle or weakening well before their advertised limit.
Does RAG remove the context-window limit?
No. RAG chooses which passages to send so you do not have to send everything, but the retrieved text still occupies the window. It makes a limited window go much further; it does not make it unlimited.
Is ChatGPT's memory the same as its context window?
No. Memory stores selected facts across conversations; the context window is what the model can see in the current request. A saved memory only affects an answer once the app inserts it into the context window.
Sources
- Liu, N. F. et al. “Lost in the Middle: How Language Models Use Long Contexts.” Transactions of the ACL, 2024.
- Hsieh, C.-P. et al. “RULER: What's the Real Context Size of Your Long-Context Language Models?” 2024.
- Petrov, A. et al. “Language Model Tokenizers Introduce Unfairness Between Languages.” NeurIPS 2023.
- Vaswani, A. et al. “Attention Is All You Need.” 2017.
- Dao, T. et al. “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.” 2022.
- OpenAI: What are tokens and how to count them?
- OpenAI Tokenizer
- OWASP: LLM01 Prompt Injection
- Current context-window sizes: OpenAI models, Anthropic, and Google Gemini.