The Context Window: The Central Constraint
Instructions, history, tools, user input, output, and sometimes reasoning all draw on finite per-call capacity, so adding context creates real tradeoffs.

A context window is the finite token capacity available to one model invocation. Everything the call needs must fit within the limits enforced by that model and runtime.
That makes context a resource-allocation problem. Instructions, conversation history, retrieved material, tool definitions, and the current user input consume capacity before the answer is complete. Output needs room too. Some models also account for internal reasoning tokens within the request’s limits.
The exact accounting is provider- and model-specific. The engineering consequence is portable: adding one kind of context leaves less capacity or headroom for something else.
Illustrative 100-token window
One call / multiple consumers130 requested / 100 available
- trim or compact input
- combine input and output reductions
- use a larger supported window
- reject or split the request
One limit, several consumers
The phrase context window is often used as though it means “how much I can paste.” That captures only the input side.
A call may need capacity for:
- platform, developer, or system instructions;
- earlier conversation turns or a compacted summary;
- retrieved documents and application data;
- tool names, descriptions, schemas, calls, and results;
- the latest user content;
- the model’s generated output;
- reasoning tokens for models and APIs that account for them.
Providers expose these limits differently. A platform may advertise one total context size plus a smaller maximum output, reserve capacity internally, or reject combinations that fit one published number but violate another constraint. Treat the target model’s current documentation and measured API behavior as the contract.
The shared-budget model is still useful because it predicts pressure. If an application adds a long tool catalog and replays more history, it should expect less flexibility elsewhere unless capacity increases or another component shrinks.
Output needs headroom
Suppose a request is close to its input limit and asks for a detailed migration plan. Even if the input is accepted, the system may not have enough permitted output capacity for the requested answer.
This is why robust applications budget before the call:
available request capacity
- required instructions
- selected history
- retrieved evidence
- tool definitions
- current user input
- output and reasoning allowance
= remaining margin
That expression is a planning checklist, not a universal billing formula. Token categories, caching, reasoning accounting, and output caps differ across implementations.
What happens near the limit
When a request is too large, a system must do something concrete. Depending on the application and provider, it may:
- reject the request with an error;
- truncate some input;
- reduce the allowed output;
- summarize or compact older state;
- retrieve a smaller selection of evidence;
- expose fewer tools;
- split the work into multiple calls.
There is no universal drop order. One framework may keep recent messages and summarize older ones. Another may require the application to trim explicitly. A provider can change what it supports across APIs and model versions.
Do not build correctness around an assumed policy such as “the oldest messages always disappear first” unless the specific runtime documents and tests that behavior. If a fact is load-bearing, make its retention deliberate.
Compaction changes the evidence
Compaction replaces some original context with a shorter representation, often a summary. That can preserve goals, decisions, and recent continuity while freeing space. It cannot promise that every detail, qualifier, identifier, or contradiction survives unchanged.
The causal tradeoff is unavoidable:
- The original history is too large to keep in full.
- A system selects and condenses information.
- The next call sees the condensed representation instead of all original tokens.
- Anything omitted or distorted is no longer available in the same form.
Compaction is useful precisely because it is selective. Treating it as lossless memory hides the decision that matters: what did the compactor preserve, and what did it discard?
Fitting is not the same as using well
A larger window relaxes a capacity constraint. It does not automatically make a model more intelligent, make every included fact equally salient, or tell the system which evidence deserves priority.
Long inputs can contain duplicated instructions, conflicting versions, irrelevant logs, and a critical line buried among thousands of tokens. Even when all of that fits, the application still has to choose and arrange context deliberately. The accepted evidence for this chapter does not establish one universal quality curve for long context, so claims about exact degradation rates should remain benchmark- and model-specific.
A prediction you can make
Imagine one request containing a long policy document, 20 chat turns, three tool schemas, several retrieved logs, and a demand for a detailed answer.
You do not need to know the provider’s exact implementation to predict the pressure:
- Every model-visible component occupies finite capacity.
- The requested answer needs output headroom.
- If the combined requirements exceed the enforced limits, the system must reject, shorten, truncate, compact, select, or split something.
- Even if the request fits, relevance and layout still affect whether the useful evidence is easy to apply.
The context window is not permanent memory and not a measure of intelligence. It is the bounded workspace for one invocation. Managing that workspace at scale is a larger topic; for a single call, the essential habit is to account for every consumer before the limit makes the decision for you.
References
- Conversation stateOpenAI
- Prompting best practicesAnthropic
- LLM Powered Autonomous AgentsLilian Weng
- How the agent loop worksAnthropic