Context Engineering: Managing the Budget Deliberately
Context engineering is the application-level discipline of selecting, representing, ordering, refreshing, and budgeting the information and capabilities a model receives for a particular task.

A model call does not decide its own complete information environment. Before generation begins, a surrounding system has already made consequential choices: which instructions to apply, which history to replay, which tools to expose, which memories or documents to retrieve, what to summarize, what to omit, and how much room to reserve for the result.
Context engineering is the discipline of making those choices deliberately over time.
The term is still used in overlapping ways. Here it means application-level policy for selecting, representing, ordering, refreshing, and budgeting what a model receives for a task. It is broader than polishing one instruction, and narrower than all of AI systems engineering.
One evolving task
Policy runs around every model call- Application / 01Observe
Budget pressure, task state, evidence needs, quality signals, and freshness.
- Application / 02Choose levers
Include, exclude, summarize, retrieve, cache, replace, or defer.
- Application / 03Assemble
Preserve authority, ordering, provenance, boundaries, and output headroom.
- Model / 04Generate
Process this particular context and produce the next response or tool request.
- Application / 05Check and update
Validate outcomes, record new state, expose loss, and recalibrate policy.
Current context is an assembled representation
For one inference, current context may be composed from several sources:
- platform, system, or developer instructions;
- the latest user input;
- selected conversation history;
- examples;
- tool names, descriptions, and schemas;
- tool requests and results;
- retrieved documents or database records;
- application state and structured metadata;
- summaries of earlier work;
- information retrieved from persistent memory;
- generated plans, scratch state, or other intermediate artifacts.
These items can all influence a call or consume its budget, but they are not the same thing. A tool definition describes an available operation. A tool result records an observation. A stored preference is durable application data. A retrieved passage is a selected source candidate. A summary is a transformed representation of prior material.
The application may also reserve capacity for generated output and, under some API contracts, reasoning-related work. That headroom shares the finite window or request budget, but it is not another information source assembled into the input.
Calling all of them “the prompt” can hide who put each item there and what reliability contract it carries.
Three layers own different responsibilities
The model
The model processes the context represented to this invocation and generates the next output. It can use patterns in its parameters and information in the current context. It does not independently inspect every external store or recover application state that was never supplied.
The application or agent runtime
The surrounding software may decide:
- which instructions and authority levels apply;
- which user, task, and conversation state belongs to this call;
- which history to retain, summarize, or remove;
- which tools are relevant and permitted;
- which documents or memories to retrieve;
- which tool results are current enough to keep;
- how to order and label information;
- what input, output, cost, latency, and iteration budgets to enforce;
- what to validate before accepting the result.
Some providers implement portions of this runtime, such as hosted tools, conversation objects, or server-side compaction. That moves the implementation boundary; it does not make the policy disappear.
External systems
Databases, search indexes, document stores, logs, memory stores, source-control systems, and tools can retain information beyond one inference. They can return current or historical state when asked.
Storage alone does not make information available to the model. The system must select or retrieve it, transform it if necessary, and place an appropriate representation in current context.
Prompt engineering and context engineering overlap
Prompt engineering often focuses on how instructions, examples, or requested output formats are written. Better wording can clarify a task and reduce ambiguity.
Context engineering asks a wider runtime question: what complete information environment should this call receive?
That can include prompt-writing decisions, but also:
- selecting one current policy instead of appending five versions;
- routing only the tools relevant to the task;
- retrieving the exact source passage a claim requires;
- preserving a security constraint outside a lossy summary;
- replacing stale application state;
- ordering content while preserving authority metadata;
- reserving enough output capacity;
- logging what compaction removed.
The boundary is not a settled industry standard. The useful distinction is operational: rewriting an instruction cannot recover evidence the application failed to retrieve, remove a stale tool result the runtime keeps replaying, or create persistent storage.
The levers are different operations
An explicit context policy can use several levers.
| Lever | Decision | Typical benefit | Risk to manage |
|---|---|---|---|
| Include | What must enter this call? | Required instructions, state, and evidence are available | Unnecessary inclusion adds noise, cost, and conflict |
| Exclude | What should stay out or leave? | Reduces distraction and budget pressure | A decisive constraint or source is omitted |
| Summarize | What can become a shorter representation? | Preserves selected long-horizon state | Detail is lost, distorted, or made stale |
| Retrieve | What exact external material should be fetched now? | Supplies task-specific or fresh evidence | Recall, ranking, permissions, freshness, or reinjection fails |
| Cache | Which repeated prefix can be processed more efficiently? | Reduces eligible repetition cost or latency | Cached context remains stale, irrelevant, or too large |
| Replace | Which state has a newer authoritative version? | Keeps current state clear | Useful history is destroyed instead of marked historical |
| Defer or split | What should be handled in another call or deterministic step? | Gives one invocation a focused job | Cross-step state or verification is lost |
Selection chooses which information. Compression chooses which representation. Ordering chooses where and under what authority it appears. Prioritization chooses what receives scarce budget when everything cannot remain. Caching changes repeated processing economics. Treating those as one generic “add context” step makes failures difficult to locate.
The control loop
Context policy should adapt as the task changes.
1. Observe
Collect signals from the real workload:
- current objective and unresolved blockers;
- input growth and remaining output headroom;
- cache reads, writes, and misses;
- stale, duplicated, or conflicting state;
- retrieval confidence and source freshness;
- quality regressions at different context compositions;
- required evidence that is no longer available.
2. Choose
Apply task-specific rules. Decide what must remain exact, what can be summarized, what should be retrieved, which tools are needed, and what can be excluded.
3. Assemble
Construct the actual request. Preserve instruction authority, source provenance, tenant or task boundaries, chronology where it matters, and a clear distinction between current and historical state. Order cacheable and volatile material according to the target API without sacrificing semantic correctness.
4. Generate
Invoke the model with that bounded representation. The response may be text, structured output, a tool request, a clarification, or a failure.
5. Check and recalibrate
Validate the result against the task’s requirements and evidence. Record new observations, replace superseded state, expose what was compacted, and adjust the next call. Do not append every output merely because it is new.
This is a control loop because context needs evolve. The file that mattered during diagnosis may become irrelevant after the fix. A tool result can become stale after deployment. A retrieved document can be superseded. A useful summary can stop being sufficient when the task turns to exact evidence.
Different tasks need different context
There is no universal best context architecture.
Short isolated question
A definition or small calculation may need the current question, concise instructions, and little or no prior history. Replaying an unrelated conversation can make the task less clear.
Coding assistant
The active objective, repository instructions, relevant files, current diff, latest failing test, and selected tools may matter. Duplicated logs and resolved false leads may not. Exact API documentation can be retrieved when the active change reaches that boundary.
Support assistant
Current customer state, the applicable policy version, relevant recent interaction, and permitted actions may matter. Another customer’s history, stale account data, and every earlier support exchange should not enter merely because storage exists.
Research workflow
Source provenance, retrieval coverage, publication dates, contradictory evidence, and claims still needing support may dominate. A smooth narrative summary without recoverable citations is a poor substitute.
The right policy follows from what the current task must know, prove, and protect.
Freshness and authority beat accumulation
Context can become wrong while remaining internally coherent. An old tool result may accurately describe a system that has since changed. A stored preference may have been revoked. A retrieved policy may have a newer revision. A summary may preserve a decision that was later reversed.
A mature policy records enough metadata to ask:
- Where did this information come from?
- When was it observed or published?
- Is it current for this task?
- Has a more authoritative source superseded it?
- Should the old value remain as history, or leave active context?
Freshness checks often require consulting the current source. Replaying existing context is not a freshness strategy.
More context can make a worse task representation
“Include it in case it helps” sounds cautious, but every addition has consequences:
- irrelevant material competes with current evidence;
- contradictions become harder to resolve;
- duplicated text overweights one source;
- old state can appear current;
- important evidence can be buried;
- token cost and latency can rise;
- less capacity remains for output or reasoning under some APIs;
- a large tool catalog can complicate capability selection.
This does not mean shorter is always better. It means usefulness is task-dependent and must be evaluated. The goal is not the smallest prompt. It is the smallest adequate, current, well-attributed representation that supports the task’s required outcome and checks.
A minimum operating policy
A long-running system should be able to answer:
- Which constraints and objectives are protected from silent loss?
- Which state is current, which is historical, and which is superseded?
- What was selected, summarized, retrieved, cached, or dropped for this call?
- Can exact evidence be recovered when a summary is insufficient?
- Which tools and permissions are exposed for this task?
- How are cost, latency, quality, and evidence availability measured together?
- What failure causes the system to retrieve again, ask for clarification, escalate, or stop?
Those answers will differ by product. Their existence is the durable practice.
Context engineering does not eliminate hallucination, guarantee long-context performance, or turn a model into persistent memory. It makes the surrounding system’s choices explicit: observe what the task needs, compose a bounded and current representation, check what happened, and revise the policy before the next call.
References
- Building effective agentsAnthropic
- Conversation stateOpenAI
- CompactionOpenAI
- Prompt cachingAnthropic
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksNeurIPS, 2020