Mental Simulation: Thinking Like the Machine
Mental simulation turns an unfamiliar AI feature or failure into a testable system hypothesis by tracing payload, prediction, tools, state, context, orchestration, production controls, and trust boundaries.

An “AI project manager” assigns tasks, summarizes meetings, updates tickets, and asks specialist agents to draft implementation plans. One day it approves work that should have waited for a human.
“The AI decided to be careless” sounds like an explanation. It does not identify a mechanism, an owner, or a check.
A useful diagnosis reconstructs the system:
- What payload and policy reached the relevant model call?
- Which model output was a suggestion, and which host action changed state?
- What meeting notes, tickets, or retrieved documents entered context?
- Which tools and subagents ran, and what did they return?
- Where should permission or approval have stopped the action?
- Which trace, test, or metric can show what actually happened?
This is mental simulation: narrating likely system behavior through explicit mechanisms, then testing the narration against observable evidence. It is not mind-reading.
Start with the complete path
For an unfamiliar feature or failure, make one pass through the system in causal order.
| Layer | Simulation question | Evidence that could answer it |
|---|---|---|
| Product and payload | What user input, instructions, history, tool schemas, and injected data formed the request? | Request assembly logs, configuration, or a privacy-safe prompt trace |
| Model call | Which model version and runtime settings produced which visible output? | Response envelope, model version, token and timing metadata |
| Orchestration | What host rule routed, repeated, delegated, accepted, or stopped the call? | Workflow definition, state transitions, route and terminal events |
| Tools and side effects | Which action did the model propose, who authorized it, and what executed? | Tool request, permission decision, operation ID, and system-of-record status |
| Memory and retrieval | What persisted outside the call, what was selected, and what entered current context? | State reads, retrieval results, source versions, and provenance |
| Context budget | What was included, summarized, cached, omitted, or displaced? | Token counts, compaction events, cache metadata, and context policy |
| Delegation | Which worker received which context and returned which result? | Delegation contract, artifacts, trace events, and validation status |
| Reliability | What behavior was expected, measured, retried, degraded, or rolled back? | Tests, task evaluations, SLIs, SLOs, alerts, release records, and incident logs |
| Security | Where did untrusted data and consequential authority meet? | Permission scope, action gates, approvals, sandbox policy, and audit events |
The table is a diagnostic pass, not a claim that every product uses one architecture or vocabulary. A simple text-completion feature may have no tools or subagents. A deterministic router may make no model call. Skip layers that do not exist; do not invent them to complete the story.
Restore the load-bearing mechanisms
Across this Series, a few distinctions do most of the explanatory work.
Input is assembled
The visible user message can be one part of a larger payload. System instructions, history, examples, tool schemas, retrieved material, and application state may also influence output and consume context.
When behavior changes unexpectedly, ask whether the represented task changed before assuming the model changed.
Model calls are bounded prediction
A model produces output from current input and learned parameters. Sampling can introduce variation. Runtime conversation continuity, durable memory, tool execution, permissions, and retries are provided by systems around the call.
Do not attribute every product action to “the AI.” Identify the model output and the host decision separately.
Agent behavior is orchestration
An agent loop can alternate model calls, tool actions, observations, and host decisions. Multi-agent systems compose more of these boundaries; they do not introduce a new kind of intelligence merely by assigning roles.
Ask who owns the next decision, what context they receive, what authority they hold, and what stops the loop.
Knowledge has a path
Parameter knowledge, retrieved documents, application state, and tool results reach a call through different mechanisms. Retrieval returns candidates, not guaranteed answers. A citation can exist without supporting the claim.
Ask where a fact came from, whether the source is current, and what independent check fits its consequence.
Context is finite and selected
Larger context can fit more information while still using it unevenly. History replay can raise cost. Compaction preserves selected information and loses detail. Caching can change processing economics without changing what the request means.
Ask what was omitted or transformed, not only whether the request fit.
Reliability and security are system properties
Prompt wording can guide behavior. Surrounding code must still validate outputs, constrain tools, apply permissions, control retries, enforce stopping, observe failures, and require human approval where appropriate.
Ask which invariant is behavioral and which component can actually enforce it.
Turn observations into hypotheses
Suppose the project manager approved a ticket unexpectedly. Several mechanism classes could produce that symptom:
- The approval policy was absent, stale, or omitted from current context.
- A retrieved ticket or meeting note contained instruction-like text that influenced the model.
- A route sent the task to a worker with broader permissions than intended.
- A worker returned
approvedin a valid schema but with the wrong semantic value. - Host code treated the model’s proposal as authorization instead of checking policy.
- A retry duplicated a side-effecting update after an ambiguous first result.
- The user interface displayed “approved” for a state that was actually “awaiting approval.”
- A release changed the workflow or permission configuration.
These are hypotheses, not interchangeable explanations. Each predicts different evidence.
| Hypothesis | Discriminating probe |
|---|---|
| Policy omitted from context | Inspect the assembled payload and policy version for the affected call |
| Untrusted content influenced behavior | Trace retrieved sources and identify instruction-like data crossing the boundary |
| Worker had excessive authority | Inspect route, tool allowlist, credentials, and approval policy for that worker |
| Structured value was semantically wrong | Compare the result with authoritative records and acceptance rules |
| Proposal bypassed authorization | Trace the model output through host validation to the tool request |
| Retry duplicated an action | Query the system of record by durable operation identifiers |
| UI mislabeled state | Compare stored workflow state with rendering logic |
| Release regression | Correlate failure onset with version and configuration changes; reproduce or roll back |
A good probe can disconfirm the favorite story. Looking only for evidence that supports one explanation turns mental simulation into confirmation bias.
State what you know, infer, and cannot see
Use three explicit categories.
I know
Reserve this for observable evidence: a request field, log event, tool result, state transition, test outcome, source passage, or measured symptom.
I know the payment tool returned operation ID P-481 and the workflow stored status `approved`.
I infer
Tie the inference to a mechanism and make it falsifiable.
I infer that host authorization was bypassed because the stored transition follows the model proposal with no approval event in the trace.
I would verify
Name the missing evidence or experiment.
I would inspect the workflow version and permission policy active for P-481, then reproduce the case with the tool replaced by a non-side-effecting test double.
This habit prevents fluent explanation from impersonating certainty. Missing traces are not permission to fill the gap with a psychological story about the model.
Predict failure before release
Mental simulation is useful during design as well as incident response.
For a proposed workflow, walk through a non-happy path:
input is ambiguous
-> router chooses the wrong specialist
-> specialist lacks one decisive document
-> valid structured result contains a wrong conclusion
-> manager accepts it without source checking
-> tool request crosses a permission boundary
-> network response is lost after the side effect
-> retry repeats the action
-> final status looks successful
At each arrow, ask:
- What information crosses this boundary?
- What can be lost, stale, duplicated, or misinterpreted?
- What authority moves or remains?
- Which component can detect the failure?
- What should the workflow do next?
- Which event proves that it stopped safely?
The exercise should lead to concrete design changes: a narrower route, an evidence field, a deterministic check, an operation ID, an approval gate, a checkpoint, a terminal failure state, or a monitoring signal. It should not lead automatically to another agent.
Prefer the earliest discriminating check
When a final answer is wrong, begin at the first mechanism that could explain it and use the cheapest evidence that distinguishes nearby possibilities.
If a retrieved policy answer is stale, inspect the retrieval result before rewriting the final prompt. If a tool action is duplicated, inspect operation status and retry logic before changing the model. If a worker’s conclusion is unsupported, inspect its returned evidence before adding a reviewer.
This keeps diagnosis tied to ownership:
- payload defects belong to assembly and context policy;
- retrieval defects belong to query, index, source, and selection behavior;
- malformed interfaces belong to schema and parsing controls;
- unsupported claims belong to evidence and verification policy;
- unauthorized actions belong to application and tool permission boundaries;
- runaway execution belongs to orchestration and stopping rules;
- production regressions belong to evaluation, monitoring, release, and recovery systems.
Model behavior can still be the relevant variable. The point is to earn that conclusion by excluding surrounding mechanisms that the model neither owns nor observes.
The checklist has limits
Mental simulation is a decision aid. It cannot recover information that was never recorded, reveal hidden model state, prove a generated rationale is causally faithful, or replace empirical evaluation.
Provider APIs, frameworks, model versions, runtime boundaries, and logging policies differ. A UI may hide orchestration details. Privacy controls may intentionally withhold prompt contents. Hosted services may own pieces of the workflow that application developers cannot directly inspect.
In those cases, state the boundary. A sound diagnosis can end with:
The visible evidence is consistent with either a retrieval omission or a downstream
validation failure. The product does not expose the source trace needed to distinguish them.
That is more useful than choosing the more dramatic story.
A compact simulation template
When you meet a new AI feature or failure, ask:
- Payload: What information and instructions entered the relevant call?
- Prediction: Which model produced which output under what settings?
- Control: What code routed, repeated, accepted, or stopped it?
- Tools: What was proposed, authorized, executed, and returned?
- State: What persisted, who could read or write it, and what entered context?
- Knowledge: Which evidence path supplied the claim, and was it current?
- Context: What was selected, compressed, cached, or omitted?
- Delegation: Which boundaries hid or transformed work?
- Reliability: What test, metric, target, recovery, or human control applied?
- Security: Where did untrusted data meet permissions or consequential action?
- Uncertainty: What is observed, inferred, and still unknown?
- Validation: Which probe could prove the diagnosis wrong?
Thinking like the machine does not mean pretending to be inside it. It means replacing the single magical box with a system you can question: explicit inputs, bounded prediction, host-controlled actions, selected context, external state, observable events, enforceable permissions, and hypotheses that remain accountable to evidence.
References
- Orchestration and handoffsOpenAI
- Integrations and observabilityOpenAI
- Monitoring Distributed SystemsGoogle SRE
- 2025 Top 10 Risk & Mitigations for LLMs and Gen AI AppsOWASP GenAI Security Project