Statelessness: No Memory Between Calls
A model call does not privately carry a conversation forward; continuity exists because an application, API, or framework supplies the relevant state again.

A chat can remember a project name from ten messages ago. That observation is real. The usual explanation, however, is not that the model call kept a private memory after it ended.
The continuity lives in a state-carrying layer around the call. An application can resend earlier messages. An API can restore a stored conversation object. A framework can load a session and assemble the next request. In each case, prior information becomes available because something supplied or restored it for the current call.
This boundary matters whenever you debug a forgotten instruction, test a correction, estimate token use, or decide where product memory should live.
One call, one supplied context
A deployed language model generates from its parameters and the context made available for this invocation. The previous call does not leave an ordinary conversational afterimage inside the next one.
Consider the same follow-up in two illustrative requests:
Request A
history: "Service Atlas times out after 30 seconds."
user: "What should I inspect first?"
Request B
history: [not supplied]
user: "What should I inspect first?"
Request A gives the model a service name and a symptom. Request B does not. The second call may still produce a plausible troubleshooting answer because the model has learned broad software patterns, but it has no request-local basis for knowing that the missing subject was Atlas or that the timeout was 30 seconds.
That is what stateless at the call boundary means here: the call’s runtime behavior depends on the context supplied to that call, not on an undocumented private transcript carried over from an earlier invocation.
It does not mean a chat product must forget everything. Products can persist plenty of state. The useful question is where that state is stored and how it returns to the next call.
How continuity is constructed
The most direct implementation keeps a message list in the client or application:
append user message to history
send history in a model request
append returned assistant message to history
repeat
The next request includes the earlier turns, so the model can condition its response on them. This is often called history replay.
Other systems hide more of that plumbing:
- An API-managed conversation object can store messages, tool calls, and tool results, then associate them with a later request.
- A response identifier can let a provider thread earlier state into a new response.
- A framework session can reload history, summaries, preferences, or application data before it runs the model.
- An application can retrieve selected memories from its own database instead of replaying the full transcript.
These mechanisms differ operationally. They share one causal shape:
- Some layer stores or identifies prior state.
- That layer restores relevant information for a later invocation.
- The restored information affects the effective context.
- The model can now produce a response that appears continuous.
Managed state can make replay invisible to the end user and even to some application code. It does not turn conversation history into model parameters.
Instructions also need a carrier
Suppose one request says:
Call this project Atlas, never Apollo.
A later response can follow that correction when the correction remains in the supplied history or another state mechanism restores it. Start an unrelated call without the correction, and there is no general reason to expect the instruction to remain active.
Some APIs also distinguish request-scoped instructions from conversation state. An instruction field can govern one response without automatically being carried through a response identifier into the next. Exact behavior is provider-specific, so inspect the API contract rather than assuming that every field persists with the thread.
Adaptation is not durable learning
Within one conversation, a model can adjust quickly to a name, formatting rule, or correction. The causal chain is usually straightforward:
- The correction enters the current context.
- Later requests carry that correction forward.
- Generation is conditioned on the corrected information.
- Responses change while that information remains available.
This is in-context adaptation. It is different from training or fine-tuning, which changes durable model state through an explicit optimization process. A successful follow-up does not demonstrate that the deployed base model rewrote its parameters or will preserve the correction in unrelated sessions.
The distinction also avoids an overclaim in the other direction. A provider may retain data, offer personalization, update a separate memory store, or later use data in a training process. Those are product architecture and policy questions. They do not change the narrower diagnosis of an ordinary model call: trace continuity to an identifiable state-carrying mechanism.
The fresh-request test
When a system seems to remember, first introduce an uncommon canary value such as 37.4 seconds, then compare two controlled requests:
- one with the relevant history, conversation reference, or restored session;
- one with the same latest message but none of that prior state.
If the second request says Use the same timeout as before, the model may guess a timeout. It cannot ground that answer in the missing earlier value. If a system returns the uncommon canary correctly and reliably, trace where that value is being carried: replayed messages, a conversation object, retrieved memory, an application database, or another supplied source.
“The server remembers our conversation.”
A server or application may store the conversation. That is not the same claim as a model call carrying private memory between invocations. Name the component that stores state, then inspect how it is restored.
This habit turns “the AI remembered” from a vague impression into a testable systems statement. It also prepares a practical consequence: every piece of history restored for continuity occupies space in the next request and may contribute to its cost.
References
- Text generationOpenAI
- Conversation stateOpenAI
- Create a MessageAnthropic