Anatomy of a Request and Response
A model call is a structured exchange: applications assemble instructions, content, controls, and tools, while responses separate generated output from protocol metadata.

A model call is often described as text in, text out. The useful reality is more structured.
Before the call, an application builds a request envelope. It chooses a model, supplies instructions and content, sets generation controls, and may declare tools. After the call, the API returns a response envelope that can contain generated text, usage data, stop information, identifiers, or a request to use a tool.
The exact JSON changes across providers and API versions. The conceptual slots are more durable than their field names.
A conceptual request
Here is pseudo-JSON for an incident-triage request:
{
"model": "selected-model",
"instructions": "Answer as a concise incident triage assistant.",
"messages": [
{
"role": "user",
"content": "The checkout service is timing out."
}
],
"tools": [
{
"name": "search_logs",
"description": "Search production logs",
"input_schema": {
"service": "string",
"since": "string"
}
}
],
"generation": {
"temperature": 0.2,
"max_output_tokens": 300
}
}
This is not a copy-paste contract for any provider. It is a map of four different concerns:
- Routing and serving: a model selector tells the platform which deployed model should handle the call.
- Model-visible context: instructions and messages supply content that can shape generation, with authority rules that depend on the platform.
- Capability descriptions: tool definitions tell the model which operations may be requested and what arguments they accept.
- Generation controls: temperature and output limits influence decoding or runtime behavior without necessarily becoming natural-language prompt text.
That last distinction matters. The request envelope is everything sent through the API, but not every transport field is converted into literal text for the model. A platform may validate some fields, use some to select infrastructure, translate some into model-specific representations, and apply others in the decoding process.
So “what crossed the wire?” and “what became model-visible context?” are related questions, not identical ones.
Authority is part of the structure
Instructions are not merely paragraphs concatenated in arbitrary order. Modern APIs commonly distinguish platform, developer or system, user, assistant, and tool-originated content. The names and exact hierarchy vary, but the reason for roles is stable: the runtime needs to preserve who supplied a piece of content and how much instruction authority it has.
A user message can contain data to analyze without gaining the same authority as a platform or developer instruction. A tool result can provide evidence without automatically becoming a command. Flattening every field into “the prompt” can hide this distinction.
The safe portable claim is not that all providers implement one universal role table. It is that role and channel metadata can change how content is interpreted, so request debugging must preserve it.
A response is more than visible text
A final-answer response might look conceptually like this:
{
"output": [
{
"type": "message",
"content": "Check the upstream timeout and connection-pool saturation first."
}
],
"stop_reason": "complete",
"usage": {
"input_tokens": 420,
"output_tokens": 18
},
"request_id": "..."
}
The assistant message is what a chat interface may render. The surrounding fields explain how the call ended, how much work was counted, and which request should be located in logs.
Another response may contain no final answer yet:
{
"output": [
{
"type": "tool_call",
"name": "search_logs",
"arguments": {
"service": "checkout",
"since": "15m"
}
}
],
"stop_reason": "tool_requested"
}
For a client-managed tool, this means the model produced a structured request. The application must validate the arguments, decide whether execution is allowed, run the external operation, and return the result in another protocol step. The language model did not execute the log search merely by emitting JSON.
Some providers also offer hosted tools that run on provider infrastructure. In that case the platform can execute the operation, but the distinction still matters: model generation proposed or selected a tool action; external runtime code performed it.
The stable map across APIs
Provider schemas differ, but most debugging questions land in a small conceptual map:
| Question | Request or response surface |
|---|---|
| Which model and runtime handle this? | Model selector and platform configuration |
| What information can shape this generation? | Instructions, messages, history, supplied context |
| What external actions are available? | Tool definitions and execution policy |
| How is generation constrained? | Decoding parameters, output limits, response format |
| What did the model produce? | Text, structured output, or tool-call items |
| Why did the turn stop? | Stop reason or lifecycle status |
| What did the call consume? | Usage and accounting metadata |
| How do we trace it? | Request, response, conversation, or event identifiers |
An API may combine slots, omit them, or expose additional ones. The map is for reasoning, not schema validation.
A better debugging record
If a response is surprising, saving only the visible prompt and visible answer discards much of the evidence. A useful record includes:
- the model and dated API version where available;
- all instruction channels and message roles;
- the relevant history and injected context;
- tool definitions and tool-choice controls;
- generation parameters and output constraints;
- the complete response items, stop reason, and usage metadata.
“My message is what the model receives.”
Your message is one part of a request assembled by an application and interpreted by a platform. It can be surrounded by instructions, history, tools, and controls that are invisible in the chat box.
Request anatomy names the containers. The next question is more revealing: how much content did the application place inside them before the call began?
References
- Text generationOpenAI
- Function callingOpenAI
- Tool use with ClaudeAnthropic
- Create a MessageAnthropic