ReAct: Interleaving Reasoning and Action

ReAct names a reason-act-observe pattern in which external results can redirect the next decision, without requiring every system to expose reasoning or use the same control strategy.

  • Explainer
  • 6 min read
Illustration of the ReAct pattern interleaving reasoning thoughts with external actions and observations.

An agent can propose an action, receive a result, and change direction. ReAct gives that rhythm a memorable vocabulary: Thought, Action, Observation.

The vocabulary is useful because it keeps evidence inside the decision process. A system reasons from its current state, acts through an external operation, observes what happened, and reasons again from the updated state.

It is not a second kind of agent loop. It is one influential way to describe and prompt the reason-act-observe pattern inside a loop.

What ReAct originally names

The ReAct paper introduced an approach that interleaves language-model reasoning traces with task-specific actions. In its examples, a model emits a thought, selects an action such as searching or looking up information, receives an observation from an environment, and continues until it reaches an answer.

The causal advantage is not the capitalization of three labels. It is the interruption of unsupported continuation:

reason from current context
-> request an action
-> receive external evidence
-> update the next decision

A model that only elaborates its initial assumptions can produce an increasingly coherent mistake. An observation can introduce a fact the initial context did not contain: zero search results, a failed test, an unexpected file, or a permission denial. The next invocation can then respond to that evidence.

Map the labels onto the loop

Modern tool-using systems may not use literal ReAct fields, but the labels map cleanly onto the architecture already established.

ReAct label Operational interpretation Responsibility boundary
Thought Reason over the currently supplied task state and choose a possible next step. Model behavior, represented for teaching; it may be hidden, summarized, or expressed only through the next output.
Action Propose a tool request or other environment-facing operation. The model proposes; host or provider runtime decides what executes.
Observation Return evidence from the environment after the action. Tool/runtime produces a result; host reinjects relevant evidence.
Repeat or finish Continue from the updated context or produce a final response. Model output influences the path; host policy controls continuation and stopping.

This mapping prevents two common collapses. An Action label does not mean the model itself executed code. An Observation is not merely narration; it is information returned from outside the model invocation and made available to a later decision.

A ReAct-like trace

Suppose a user asks a coding agent to fix a failing test. A compact teaching trace might look like this:

Step Teaching label Event
1 Current state The task is known, but no fresh test result is present.
2 Thought More evidence is needed before choosing a file to edit.
3 Action The model requests run_tests.
4 Observation The host returns one failed assertion and a stack trace.
5 Thought The stack trace points to the configuration parser.
6 Action The model requests read_file for that source path.
7 Observation The host returns the parser implementation.
8 Next step The model can propose an edit grounded in the observed failure and code.

The visible Thought rows are pedagogical summaries. A real product may expose a brief rationale, show no reasoning text, or represent decisions only as structured tool requests. The reason-act-observe dynamic does not depend on giving users raw chain-of-thought.

Observation is the course-correction step

Consider a search for loadLegacyConfig.

If the search returns two files, the next call can choose which file to inspect. If it returns zero files, the next call can broaden the query or ask whether the symbol was renamed. If it returns a permission error, the next call can request access or report a block.

Without a returned observation, all three worlds look the same to the next model invocation. It may continue writing as though the expected file exists because no external evidence corrected that assumption.

The mechanism is:

  1. an action changes or inspects the environment;
  2. the environment produces a result;
  3. the host places that result in later context;
  4. the next output distribution is conditioned on different evidence;
  5. the trajectory can change.

Observation improves grounding only to the extent that the evidence is accurate, relevant, fresh, and represented well. A broken tool can return wrong data. A host can truncate a result. A model can misinterpret a correct result. ReAct creates an evidence path; it does not guarantee truth.

“Thought” is a conceptual label, not an inspection port

The original ReAct method uses generated reasoning traces in its prompted trajectories. That research setup should not be turned into a universal claim about production systems.

Providers and products differ in what they expose. A system may:

  • hide internal reasoning and show only tool events and final output;
  • provide a short rationale or summary;
  • generate visible planning text;
  • use structured decisions without a field named thought;
  • follow a ReAct-like cycle while presenting no trace at all.

Visible reasoning text should therefore be read as output or product presentation, not assumed to be a complete causal transcript of model computation. ReAct can explain event sequencing without resolving the broader question of reasoning faithfulness.

ReAct does not replace orchestration

The pattern says little by itself about permissions, retries, budgets, or termination. Those remain host policies.

A model may emit an action after a tool error. The host still decides whether that action is allowed. A model may produce a final answer. The host can still require a test result before accepting completion. A model may keep requesting the same failing operation. The host needs a no-progress or iteration limit.

ReAct also is not the only valid control strategy. Systems can use fixed workflows, plan-then-execute stages, batched tools, parallel operations, evaluator loops, or human approval gates. Some preserve the reason-act-observe rhythm without using the name. Others solve a well-defined task more predictably with ordinary code and one model call.

“Every agent must show Thought, Action, Observation, so the visible thoughts reveal how it reasoned.”

ReAct is a conceptual and prompting pattern, not a required provider protocol. A system can interleave decisions, actions, and observations while hiding or differently representing reasoning. Visible reasoning is not automatically a complete account of internal computation.

Use the pattern as a diagnostic

ReAct is most useful when it helps identify a missing link:

  • Thought without Action: the system describes what it would inspect but never requests an operation.
  • Action without Observation: a tool runs, but its result never reaches the next call.
  • Observation without revision: new evidence arrives, but the system repeats a stale plan.
  • Iteration without stop policy: the cycle continues despite failure or no progress.
  • Visible reasoning without evidence: the interface narrates a convincing process, but no external event supports it.

The last case leads to a practical reading skill. Agent interfaces compress requests, tool events, model calls, and results into phrases such as “Searching files” or “Reviewed 8 files.” Those phrases can be useful, but they are product narration over the loop, not the loop itself.

References

  1. ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao et al.
  2. Building effective agentsAnthropic
  3. OpenAI Model Spec, February 12, 2025OpenAI