The Rising Cost Curve: Why Sessions Get Expensive
In a long session, each later call may reconstruct more prior state than the one before it, so per-turn input, cumulative work, latency, and context pressure can rise even when new messages stay short.

At the start of a session, a five-word question may accompany a small amount of prior state. Hours later, another five-word question may accompany instructions, a growing history, tool requests and results, summaries, retrieved evidence, and generated intermediate state.
The visible message stayed short. The call did not.
That is why cost is not necessarily flat per message. In systems that reconstruct continuity for each model invocation, later turns can carry more input than earlier turns. Per-turn cost tends to rise with that input, and cumulative session cost grows faster because earlier information may be processed repeatedly.
Directional session sketch
Payload growth, not provider pricing- Early
Stable prefix + short history
- Middle
Stable prefix + growing replay
- Late
Stable prefix + long full history
- Early
Cache begins warming + short state
- Middle
Cached prefix + compacted state
- Late
Cached prefix + selected state + retrieval
- Stable prefix
- Same prefix, cache-eligible processing
- Volatile session state
Continuity is reconstructed per call
A language model processes the context available to a particular inference. A surrounding application may maintain a conversation object, store messages, or pass a previous response identifier, but the model call still needs a representation of relevant prior state.
In a simple client-managed conversation, the pattern is visible:
turn 1 request = instructions + question 1
turn 2 request = instructions + question 1 + answer 1 + question 2
turn 3 request = instructions + question 1 + answer 1
+ question 2 + answer 2 + question 3
If each request replays all earlier content, the input to an individual call grows as the session grows. The total work across the whole session includes the same early material again and again.
Provider-managed conversation state can hide that replay from application code. It does not justify assuming the context is free. For example, OpenAI currently documents that chaining Responses with previous_response_id still bills previous input tokens in the chain. Other APIs expose different accounting and caching behavior. The usage record for the actual model and API is the contract.
Three curves matter
“The cost curve” is useful shorthand, but a long-running system should observe at least three related trends.
Input per turn
How many input tokens or provider-specific units does each invocation process? This reveals whether history, tool results, retrieved documents, or instruction catalogs are accumulating.
Cumulative session usage
How much input, output, reasoning work, and external tool activity has the workflow consumed in total? A modest increase per turn can still produce substantial cumulative growth across many calls.
Time and quality
More input can affect time to first output, total latency, and available output headroom. At the same time, stale or noisy context can reduce effective use. A cheaper request that omits decisive evidence is not an optimization; an expensive request that repeatedly sends irrelevant history is not thoroughness.
These curves do not have one provider-independent shape. They depend on message size, response size, model, tokenizer, reasoning configuration, cache hits, compaction policy, retrieval, tool activity, retries, and traffic pattern. The structural prediction is narrower: unmanaged replay creates a mechanism for growth.
Stable and volatile context grow differently
It helps to separate request content by how often it changes.
Stable prefix might include:
- system or developer instructions;
- a selected tool catalog;
- fixed examples;
- a static reference document;
- stable project conventions.
Volatile state might include:
- new conversation turns;
- latest tool output;
- current application state;
- retrieved evidence for this question;
- an evolving task summary;
- generated plans or intermediate results.
This distinction reveals why caching and compaction bend different parts of the cost path.
Caching changes repeated-prefix economics
When a provider recognizes an eligible matching prefix, prompt caching can reduce the price or processing latency associated with that repeated region under the current API contract.
Caching does not necessarily shorten the context. A 50,000-token cached prefix can remain a 50,000-token prefix for window accounting even if its input rate is discounted. The system can therefore achieve a better bill while still facing output pressure, stale information, and long-context degradation.
Cache effectiveness also depends on actual reuse. If a supposedly stable prefix changes every turn, expires before reuse, falls below an eligibility threshold, or is assembled in an incompatible order, the expected savings may not appear. Reads and writes should be observed, not inferred.
Compaction changes the evolving payload
Compaction selects and transforms volatile state so later calls do not carry the full raw history.
growing conversation and tool trace
-> preserve active constraints
-> summarize resolved narrative
-> retain current state
-> retrieve exact evidence when needed
-> drop low-value duplicates visibly
That can reduce token pressure and alter the slope of per-turn growth. It can also lose information. The cost curve and the fidelity curve therefore have to be reviewed together.
If a summary removes the exception that decides the current case, lower input usage has purchased a worse task representation. If a retrieval policy can recover the exact current clause, the application may avoid keeping every policy document in every call, but retrieval recall and freshness become part of the reliability budget.
The mechanisms compose
Consider a long product-research session.
| Request region | Policy | Effect |
|---|---|---|
| Product brief and tool schemas | Keep stable and cache where supported | Reduces repeated-prefix economics when hits occur |
| Active objective and unresolved questions | Preserve in current state | Keeps the task explicit |
| Older discussion | Summarize with decisions and rejected paths | Reduces evolving payload at a fidelity cost |
| Exact source passages | Retrieve when the active claim needs them | Avoids carrying all sources continuously, subject to retrieval quality |
| Duplicate search results | Exclude or deduplicate | Removes low-value repetition |
| Latest evidence | Replace stale versions or mark history clearly | Reduces conflict and freshness risk |
Caching, compaction, retrieval, and selection are not competing names for one trick. They operate on different parts of context construction and can fail independently.
Why lower temperature does not fix structural growth
Changing a decoding parameter can affect output selection. It does not remove replayed history from the next request, create cache hits, or decide which evidence can be summarized.
Similarly, asking the model to “be concise” may reduce generated output if it follows the instruction, but it does not by itself bound tool calls, retries, or input replay. Cost controls should be attached to the layer that creates the cost:
- context policy for input growth;
- output limits and validation for generation size;
- loop budgets for repeated calls;
- tool policy for external operations;
- cache design for repeated eligible prefixes.
Thresholds are operational heuristics
Teams often ask for one trigger such as “compact at 80% of the window.” A fixed percentage is easy to implement and sometimes useful as a starting guard, but it is not a universal law.
A good trigger depends on what the workload needs to preserve and what the target runtime reports. Signals can include:
- remaining context and output headroom;
- input tokens and their rate of growth;
- cache read and write ratios;
- time to first output and total latency;
- repeated or stale content detected in assembled requests;
- retrieval availability for exact evidence;
- task-quality regressions at different context lengths;
- the risk of losing protected constraints.
A short legal clause can be more important than thousands of narrative tokens. A coding run may need the latest test output immediately but only a summary of ten rejected hypotheses. Thresholds should be calibrated against real traces and evaluations, then revisited when the model, tokenizer, provider pricing, tool surface, or workload changes.
Cost and quality share the same policy boundary
The application or agent runtime decides which history to replay, which tools to expose, what to retrieve, when to summarize, what to discard, and how much output or iteration budget remains. Those decisions affect both economics and model behavior because they shape the same context.
The model cannot recover a dropped constraint merely because dropping it saved tokens. It also cannot know that an old tool result is stale unless the application replaces it or marks it clearly.
A useful review joins four records:
- Assembly: what entered each call and why.
- Transformation: what was cached, summarized, retrieved, replaced, or dropped.
- Accounting: input, cached input, output, reasoning, tools, retries, and latency.
- Outcome: whether the response used current evidence and met the task’s requirements.
That turns a surprising bill into a system trace. The rising curve is not a mysterious property of a long chat. It follows from repeated calls carrying reconstructed state, and it can be reshaped only by deliberate policies with visible quality tradeoffs.
References
- Conversation stateOpenAI
- Prompt cachingOpenAI
- Prompt cachingAnthropic
- Building effective agentsAnthropic