Tool Calling: How a Text Model "Does" Things
Tool use is a protocol boundary: the model proposes a structured call, external software decides whether to execute it, and a later model call receives the result.

When an interface says that a model searched files, the shortest description hides the most important boundary.
The model can generate a structured request for a search. Something else must inspect that request, decide whether it is allowed, execute search code, capture the result, and make the result available to a later model call.
That is tool calling: a protocol that connects generated output to ordinary software. The model proposes. A host or provider runtime executes. A later invocation can condition its next decision on the outcome only when the result or other execution evidence returns.
A tool begins as a description
Before a model can request a client tool, the application usually describes the available operation. Here is normalized pseudo-JSON for a file search:
{
"name": "file_search",
"description": "Find files containing a text query within an allowed scope.",
"input_schema": {
"query": "string",
"scope": "string"
}
}
The names and shape are illustrative. OpenAI, Anthropic, and other providers use different request fields, content blocks, stop signals, and schema conventions. Some frameworks translate their own tool definitions into a provider format.
The durable mechanism is that a tool description becomes part of the call’s available context or capability configuration. It tells the model what operation can be proposed and what arguments the host expects.
Nothing has run yet. A schema is a description of an interface, not an executable event.
The model emits a proposal
Given a task and the tool description, a model response might contain a normalized request like this:
{
"decision": "tool_call",
"tool": "file_search",
"arguments": {
"query": "loadLegacyConfig",
"scope": "src/**"
}
}
At this point, one fact is established: the model generated data that identifies a tool and proposes arguments.
The following facts are not yet established:
- that the named tool exists in the current runtime;
- that the arguments satisfy application policy;
- that execution began;
- that the search succeeded;
- that any returned files are relevant;
- that the model has seen a result.
Calling the structured output a tool request preserves that uncertainty. It is more precise than saying the model searched the repository.
The host crosses the execution boundary
For a client-managed tool, application code receives the request and controls what happens next. A responsible host can:
- confirm that
file_searchis registered; - parse and validate the arguments;
- constrain the requested scope;
- check permissions or ask for approval;
- execute the search implementation;
- capture output, failure, timeout, or cancellation.
Those steps are ordinary software behavior. They can be tested, logged, rate-limited, denied, retried, or placed behind a human decision.
Provider-hosted tools move the execution code to provider infrastructure. That changes which external runtime performs the operation, but it does not collapse generation and execution into one event. The model output still participates in a protocol; platform code performs the search, browser request, code execution, or other operation.
This distinction matters most for state-changing actions. A model might emit a perfectly formed request to delete files, send an email, or issue a refund. Well-formed data does not grant authority. The executing layer remains responsible for policy, permissions, target validation, and any required approval.
A result must come back as evidence
Suppose the search succeeds. The host can represent the observation in another normalized record:
{
"tool_result": {
"tool": "file_search",
"status": "ok",
"matches": [
"src/config/loadLegacyConfig.ts",
"src/bootstrap/startup.ts"
],
"match_count": 2
}
}
Execution produced evidence outside the model call. To affect the next model decision, that evidence must enter a later context package. Providers may represent the return as a tool-result message, a content block, a response item, or managed conversation state. The field names vary; reinjection is the causal step.
The next call can now propose reading two specific files because the match list is available. Without the returned result, a later call cannot distinguish a successful two-file search from a timeout, an empty result, or an operation that never ran.
Failures are observations too:
{
"tool_result": {
"tool": "file_search",
"status": "error",
"error": "scope_not_allowed"
}
}
If the application reinjects that error, a later model call can choose a narrower scope, ask for permission, or explain that it is blocked. If the application hides the error and silently calls the model again, the next output has no grounded basis for adapting to it.
Who owns each step
| Step | Primary owner | What it establishes |
|---|---|---|
| Choose which tool to request and propose arguments | Model output | A candidate next action |
| Decide which tools are exposed | Application or platform | The available action surface |
| Validate arguments and permissions | Executing host/runtime | Whether the request may proceed |
| Perform the operation | Client application, provider runtime, or another tool service | An external effect or observation |
| Record the outcome | Host/runtime | Structured evidence of success, failure, or partial completion |
| Supply the result to another call | Application, API, or framework | Model-visible evidence for continuation |
The boundaries can cross organizations. An application may call a provider, which invokes a hosted tool, which calls another service. Use the same diagnostic question at every boundary: who generated a proposal, who authorized it, who executed it, and which result reached the next decision?
Schemas constrain shape, not judgment
A strict schema can prevent missing fields or wrong primitive types. It cannot guarantee that the model selected the right tool, chose a sensible value, understood the user’s intent, or accounted for a dangerous side effect.
For example, this request can be structurally valid and operationally wrong:
{
"tool": "issue_refund",
"arguments": {
"order_id": "A-1842",
"amount": 5000
}
}
The host still needs domain checks: does the order exist, is the amount in the expected currency unit, is it within policy, and did the user authorize the action? Schema validation is one gate, not the whole safety model.
Likewise, providers may allow zero, one, several, or parallel tool requests in one model response. A product may automatically run safe read-only operations while gating writes. These are implementation choices, not properties shared by every tool-using system.
“The model called the tool, so the model executed it.”
In precise terms, the model produced a tool-call request. Client application code or provider runtime code performed the operation. The next model call can use the outcome only after the system makes that outcome available as context.
One request-execute-result cycle explains a single handoff. An agent emerges when software repeatedly assembles state, invokes the model, dispatches approved tools, reinjects observations, and applies explicit continuation and stop rules.
References
- Function callingOpenAI
- Tool use with ClaudeAnthropic
- Create a MessageAnthropic