The Agent Loop: Orchestration Around the Model

A practical agent is a controlled system around model calls; the orchestrator assembles state, dispatches tools, handles failures, and decides when to repeat or stop.

  • Explainer
  • 7 min read
Illustration of host-controlled orchestration surrounding probabilistic model calls in an agent loop.

A tool request explains one handoff. An agent loop explains how a system keeps making handoffs until it finishes, becomes blocked, or is told to stop.

The model is one component inside that loop. Surrounding software assembles the context for each invocation, interprets model output, executes approved tools, records observations, and applies continuation policy. That software is often called the host, runtime, or orchestrator.

Different frameworks divide the work differently. The normalized loop is still useful because it exposes the control flow that a product name or circular arrow can hide.

One normalized iteration

Continuation remains host-controlled
  1. Host / 01Assemble context

    Instructions, selected state, prior observations, and available tools

  2. Model / 02Generate a decision

    A candidate tool request, clarification, or final response

  3. Host / 03Interpret the output

    Validate its shape, apply policy, and choose the next branch

Tool requestedValidate, execute, capture

The host or provider runtime runs approved code and records a result or error.

Observation returns to context assembly
Final or other outputApply stop policy

The host may finish, ask for clarification, retry, reject, or report a blocked run.

Stopping is a policy decision
Model behavior
Proposes next actions and generates language from supplied context.
Application behavior
Owns state, permissions, budgets, retries, continuation, and termination.
Tool/runtime behavior
Performs external work and returns observable success, output, or failure.
The orchestrator controls the loop.This is a normalized control-flow skeleton. Frameworks and providers may split, combine, batch, or host stages differently, but generated proposals remain distinct from execution and stop policy.

The minimum viable control flow

One iteration begins before the model is called.

1. Assemble the current state

The host selects what this invocation can use. That package can include:

  • the user’s task and active instructions;
  • relevant conversation history or a summary;
  • prior tool requests and results;
  • files, plans, diffs, or other workspace observations;
  • tool descriptions and permissions;
  • remaining time, token, cost, or iteration budgets.

This is state assembly. It does not place durable memory inside the model. It builds the bounded context from which the next output will be generated.

State selection has causal force. If the latest test failure is included, the model can propose a change grounded in that failure. If the host supplies an older passing result instead, the model may reasonably act as though the problem no longer exists. Apparent agent confusion can therefore originate in stale or missing application state even when the model service is operating normally.

2. Invoke the model

The host sends the assembled request. The model produces a candidate next output: perhaps a tool request, a clarification question, structured data, or a final response.

That choice is probabilistic model behavior shaped by the supplied context and decoding process. The host can constrain the available tools, instructions, output schema, and budget, but it does not fully determine which valid output the model will produce.

3. Interpret and gate the output

The host parses the response and decides which branch it represents. If it contains a tool request, the host can validate its shape, check permissions, require approval, reject it, or dispatch an allowed operation.

If the output appears final, the host can accept it, validate it, request another model pass, ask the user for missing information, or stop with an error. A model-generated final answer is one input to termination policy; it does not force the surrounding application to declare success.

4. Execute outside the model

An approved client tool runs in application code or another service. A hosted tool runs in provider infrastructure. Either runtime can succeed, fail, time out, return partial data, or be cancelled.

The host captures that outcome as an observation. The observation should preserve enough structure to distinguish, for example, an empty successful search from a failed search or a search denied by policy.

5. Reinject and decide

If work should continue, the host updates its state and assembles another request containing the relevant observation. The next model invocation can now choose a step conditioned on what actually happened.

If work should stop, the host records why: success, clarification required, user cancellation, safety denial, exhausted budget, repeated lack of progress, or an unrecoverable external failure.

The circular diagram becomes real control flow only when every repeat has a condition and every exit has a reason.

The host controls; the model proposes

A useful way to inspect an agent is to label every trace line by owner.

Trace event Layer Why
Available tools: search, read, patch, test Application control The host decides what capabilities enter this run.
Request search for loadLegacyConfig Model proposal Generated output names a candidate operation and arguments.
Search returned two files Tool/runtime observation External execution produced evidence.
Include both paths and their contents in the next request Application state assembly The host selects what the next call can condition on.
Retry at most twice, then report the failure Application policy Retry count and stop behavior are enforced outside the model.
The change is complete Model-generated claim The host may still require tests or another completion check.

The phrase host control needs one qualification. The host controls the software path, not every model choice or external outcome. It can decide which tools are available and reject a request; it cannot guarantee that the model selects the best action or that a network service succeeds.

Deterministic orchestration does not mean a deterministic run

Agent loops are often described as deterministic code around probabilistic model calls. The useful part of that statement is that orchestration rules can be explicit and testable:

assemble state
invoke model
interpret output
if an approved tool is requested:
    execute it
    record the observation
    continue if policy allows
otherwise:
    apply final, clarification, retry, or failure policy

The same code path does not make the whole run repeat byte for byte. Model output can vary. Files and web pages can change. Services can time out. Parallel operations can complete in different orders. Policies can branch on all of those results.

So the precise contrast is:

  • orchestration logic is ordinary program logic that can be inspected and tested;
  • model decisions and environment results can vary within that logic;
  • agent behavior emerges from both.

Failures must become usable observations

Tool success is the easy branch. Robust loops also represent failure.

Suppose a test command times out. The runtime should report a timeout rather than fabricate a test result. The host can then choose among several policies:

  • reinject the timeout and let the model propose a narrower command;
  • retry automatically under a bounded rule;
  • ask the user before spending more time;
  • stop and report that verification is incomplete.

Retrying is not an instinct that the model owns. A model can propose another attempt, but the application determines whether another attempt is permitted, how many attempts remain, and which failures are retryable.

The same principle handles invalid model output. A host may reject malformed arguments, return a validation error as an observation, request a repaired structure, or end the run. Silently pretending the invalid request succeeded removes the evidence needed for correction.

Stop conditions are part of the architecture

A loop without an exit policy can keep spending after it stops learning anything new. Useful stop conditions include:

  • required success evidence has been observed;
  • the model produced a final response and the host accepts it;
  • the user must clarify or approve the next action;
  • a safety or permission gate rejects continuation;
  • a time, token, cost, or iteration budget is exhausted;
  • the same action or error repeats without progress;
  • an external dependency fails in a non-retryable way;
  • the user cancels the run.

Different products prioritize these conditions differently. Some stop whenever a response contains no tool request. Others run validators or deterministic checks before accepting completion. Some let a human approve state-changing operations. None of those policies is supplied automatically by the phrase “agent loop.”

“The agent autonomously decides, acts, remembers, and stops.”

The model may influence the path by proposing next actions. The surrounding system retains state, executes or delegates tools, enforces gates, handles retries, and decides whether the run continues. Treating all of that as one model behavior makes failures harder to locate.

The skeleton is portable, not universal

Real implementations can batch several tool calls, run safe reads in parallel, combine interpretation with state updates, delegate hosted tools to a provider, or place human approval between proposal and execution. A fixed workflow may choose more branches in code; a more open-ended agent may let model output choose among tools.

Those variations matter operationally. They do not erase the responsibility boundaries:

  1. some layer assembles what the model can see;
  2. a model invocation produces candidate output;
  3. some runtime decides what can execute;
  4. external work produces observations;
  5. some layer decides what returns and whether another invocation occurs.

That is the loop to look for under the framework. ReAct gives one influential vocabulary for describing its decision, action, and observation rhythm, but it is a strategy for understanding the cycle, not a requirement that every agent expose the same labels.

References

  1. Building effective agentsAnthropic
  2. Function callingOpenAI
  3. Tool use with ClaudeAnthropic
  4. Conversation stateOpenAI