AI Under the Hood

The Round Trip: What Happens When You Hit Enter

A chat message becomes one part of an application-assembled request, a model generates from that request, and the application carries the result back.

  • Explainer
  • 4 min read
Illustration of an AI chat interaction moving from a user message through application request assembly to a model response.

Press Enter in an AI chat product and the result can feel immediate: your sentence disappears into the system, and an answer begins to appear. That view hides the useful part.

The product is more than the model. It also includes an interface and application code that prepare a request, call the model, and handle what comes back. Once those parts are visible, one interaction becomes a sequence you can inspect rather than a single mysterious act.

One interaction

The application carries state across the model-call boundary
  1. Interface / 01Submit

    The user sends a visible message.

  2. Application / 02Assemble

    The host combines the message with relevant instructions, history, context, and settings.

  3. Model / 03Generate

    The model produces a continuation from the context supplied for this call.

  4. Application / 04Handle

    The host receives the response, then displays, stores, or forwards it.

For a later turnThe host must supply or link the relevant context again.
A model interaction is a staged round trip.The visible message is only one possible part of the assembled request. Continuity comes from application-managed state returning to a later call, not from a private transcript retained by the model.

The message triggers an application

Suppose you type:

What should I do next?

The interface passes that event to the product’s application code, sometimes called the host. The host decides what to send. It may assemble a request containing:

  • instructions that shape the product’s behavior;
  • selected messages from earlier in the conversation;
  • the current user message;
  • relevant application or tool context;
  • settings for this model call.

The latest message is real input, but it is not necessarily the complete input. A different host can wrap the same visible sentence in different context and produce a meaningfully different request.

Not every product exposes all of this assembly to the user. The important beginner’s model is not a universal wire format. It is the boundary: the application assembles the call; the model receives what the application supplies.

The model generates from the current request

A language model is the trained component that produces a continuation from the context available in this call. It does not fetch a finished answer and then reveal it. At a high level, it generates the response one small unit at a time, with each generated unit influencing what can follow.

Chapter 02 will unpack those units and the generation loop. For now, keep two consequences:

  1. The response is shaped by the context actually supplied, not by everything the user intended.
  2. A plausible response is a capability, not a guarantee that the content was checked against external facts.

This keeps model and product separate. The model generates. The surrounding product may add instructions, stored history, tools, retrieval, interfaces, and safety controls. Product behavior should not automatically be attributed to the model alone.

The application handles the return

Once generation finishes, the response returns to the host. The application can display the text, save it, record usage information, or pass it into other application logic.

Those are product actions. Seeing an answer in a chat history proves that the product received and rendered it. It does not prove that the model retained a private copy after the call ended.

Continuity has to cross the boundary

Now imagine that the earlier conversation included this detail:

We are debugging a failed database deployment.

Two later requests can show the same visible question, What should I do next?, while carrying different context:

Request A
selected history: "We are debugging a failed database deployment."
current message:  "What should I do next?"

Request B
selected history: [not supplied]
current message:  "What should I do next?"

Request A gives the model a specific situation to continue from. Request B does not. Different responses should be expected because the effective requests differ.

For a later turn to use earlier information, some state-carrying mechanism must make that information available again. An application can replay messages, restore stored conversation state, link a previous response, retrieve selected information, or reuse an eligible prompt prefix. These implementations differ, but none requires a hidden transcript living inside the model call.

Prompt caching is one easy case to misread. It can reuse matching request content to reduce repeated work. That is a runtime optimization over supplied context, not persistent personal memory.

A prediction you can make

When two runs behave differently, compare their round trips before calling the difference mysterious:

  1. What did the interface show?
  2. What did the application assemble?
  3. What context was available during generation?
  4. What did the application do with the response?
  5. What, if anything, was carried into the next call?

That sequence is the low-resolution map for the whole Series. The next lesson turns its most important boundaries into five claims to carry forward. Chapter 02 then begins the first close inspection: how text becomes model input and how generation proceeds from it.

References

  1. Conversation stateOpenAI
  2. Prompt cachingAnthropic
  3. Language Models are Few-Shot LearnersNeurIPS, 2020