Demos vs. Production: The Reliability Gap
A demo proves that one trace can work; production requires the surrounding system to measure representative behavior, constrain actions, expose failures, recover safely, and operate under real variability.

A support workflow solves five scripted tickets in a meeting. The answers are fast, polished, and correct. The team has shown that the system can succeed.
It has not shown how often it succeeds across real ticket types, what happens when retrieval is stale, how latency changes under load, whether a tool action can be duplicated, how an operator finds a failed subagent, or how the product responds when confidence and dependencies degrade.
A demo is a curated run. Production is the repeated operation and control of a variable system.
The environment changes the evidence
A happy-path demo usually has selected inputs, known operators, short sessions, responsive dependencies, and immediate attention from the people who built it. Production introduces distributions and change.
| Demo evidence | Production question |
|---|---|
| Five scripted tasks passed | How does the workflow behave across representative task classes and edge inputs? |
| One response was fast | What are typical and tail latencies under real traffic and dependency load? |
| No visible tool error occurred | Which tool failures, timeouts, partial results, and unknown outcomes are observable? |
| The answer looked correct | Which claims or actions can be checked against independent evidence? |
| One release worked | Can a regression be detected, contained, and rolled back? |
| The operator watched the run | Will monitoring detect the same problem when nobody is watching? |
| The prompt stayed short | How do context, token use, and cost change during long-running work? |
Demo success is useful. It establishes feasibility and can expose obvious design flaws. It is weak evidence for frequency, tails, recovery, security, and operational cost.
Reliability belongs to the whole system
Model behavior can vary with input, context, sampling, model version, and provider behavior. Better instructions can reduce some ambiguity. They cannot enforce every production invariant.
The surrounding layers own different controls:
Model
The model produces text, structured values, or tool requests from the represented input. It may follow instructions, use supplied evidence, and identify uncertainty. Its own confident output is not an enforcement boundary or proof of correctness.
Agent or application runtime
The runtime can constrain available tools, assemble state, validate formats, apply permissions, enforce budgets, control retries, require approvals, record traces, and decide when a result is accepted or rejected.
External systems
Databases, queues, tools, identity systems, sandboxes, test runners, and systems of record enforce durable state and side effects. They can provide operation identifiers, authorization decisions, transaction status, and authoritative evidence that a model response cannot create by assertion.
Operators and reviewers
People define acceptable risk, investigate exceptions, approve high-consequence actions, and make ambiguous decisions the workflow should not automate. Human involvement is a deliberate control, not evidence that a system has failed to become “mature.”
A reliable AI system is therefore not simply a reliable model. Reliability can emerge from context policy, tools, state, orchestration, validation, security boundaries, recovery, observability, and appropriate human control.
Define behavior before measuring it
Production monitoring works best when it begins with user-visible expectations.
A service level indicator is a defined measurement. A service level objective is a target for that indicator. An error budget makes the tolerated amount of unreliability explicit instead of silently demanding perfection.
Traditional signals such as latency, traffic, errors, and saturation remain relevant. Agentic workflows may also need product-specific indicators, such as:
- completed tasks that satisfy a deterministic acceptance check;
- tool-call failure and timeout rates;
- retrieval misses or unsupported-citation rates;
- human escalation and approval rates;
- workflows ending at iteration, time, token, or cost limits;
- route changes, fallback use, and degraded-mode frequency;
- side effects with unknown or conflicting status;
- cost per completed task rather than cost per isolated call.
No universal set or target fits every product. A code assistant, a document search tool, and a payment workflow have different definitions of success and different consequences of failure.
The useful control loop is:
define expected behavior
-> measure representative signals
-> compare results with targets and policy
-> act: continue, investigate, degrade, roll back, or stop
-> learn and recalibrate
More metrics do not automatically improve reliability. Signals should lead to a decision. Low-noise monitoring helps an operator distinguish user-visible symptoms from interesting but unactionable variation.
Observe the workflow, not only the final answer
If the only retained event is the final prose, an incident can be difficult to reconstruct.
Depending on privacy and risk, useful operational evidence can include:
- task and state transitions;
- model inputs and outputs that policy permits the system to retain;
- routes, handoffs, and delegated task identifiers;
- tool requests, authorization decisions, results, and errors;
- retrieved source identifiers and versions;
- validation and approval outcomes;
- retries, timeouts, and terminal states;
- latency, token usage, and cost by stage;
- release, model, prompt, tool, and policy versions.
This is workflow observability, not a demand to store hidden model reasoning. Observable inputs, outputs, state changes, and external events are more useful for diagnosis than a generated story about why the model acted.
Instrumentation must also obey data policy. Sensitive prompts, credentials, customer records, and tool outputs may need redaction, restricted access, or shorter retention. “Log everything” is not a safe observability strategy.
Evaluation must cover more than one prompt
“It worked for my example” tests existence. Production evaluation asks how the workflow behaves over representative tasks and known failure modes.
Different checks answer different questions:
| Check | Question it can answer |
|---|---|
| Deterministic unit-like validation | Does parsing, schema handling, policy code, or a pure calculation behave as specified? |
| Integration test | Does the runtime correctly connect models, tools, state, permissions, and failure handling? |
| End-to-end task evaluation | Does the complete workflow achieve the intended outcome over representative cases? |
| Output or model evaluation | Does generated behavior meet defined quality criteria, with known evaluator limitations? |
| Production monitoring | Is real operation staying within targets, and are new failure classes appearing? |
These layers complement one another. A schema test can prove format validity but not factual correctness. An end-to-end score can hide which component failed. Production monitoring can reveal drift but should not be the first place a known dangerous case is exercised.
Evaluation sets should include ordinary tasks, edge inputs, dependency failures, long-running behavior, and risk-relevant cases. The exact methodology depends on the product; the durable requirement is to make the evidence broader than a curated trace.
Recovery is more than retrying
Retries can help with transient network failures, provider errors, or a malformed response that can be regenerated safely. They can also amplify cost, latency, rate pressure, and duplicated work.
The critical distinction is between retrying computation and retrying a side effect.
Regenerate a classification
-> usually no external change has occurred
Create an order, send an email, transfer funds, delete data
-> the first attempt may have succeeded even if the response was lost
A blind second attempt can duplicate the action. Safe behavior depends on the external operation’s contract: durable operation identifiers, idempotency where supported, status lookup, transaction boundaries, or human reconciliation. The model should not guess whether an unknown side effect happened.
Recovery policy may choose to:
- retry a bounded safe operation;
- switch to a degraded capability;
- use a different dependency or route;
- restore a checkpoint;
- roll back a release or configuration;
- ask the user for missing information;
- require human approval or intervention;
- stop with a clear failure state.
An OpenAI Responses API incident in May 2026 provides a concrete, narrow example: elevated errors were linked to a recent deploy and mitigated by rollback. The lesson is not that every AI failure is a release failure. It is that production controls must include ordinary software recovery paths; prompt changes are not the answer to every incident.
Stopping behavior is part of the architecture
Agentic workflows can continue making calls when results remain ambiguous, tools keep failing, or the model repeatedly proposes another step. Production code needs explicit termination behavior.
Depending on the workflow, it may enforce:
- a maximum number of iterations or repeated actions;
- time, token, and cost budgets;
- explicit success, failure, cancelled, and needs-approval states;
- task-specific completion checks;
- thresholds for unresolved validation failures;
- human approval before the next consequential stage.
The system should distinguish “completed,” “could not complete,” and “stopped by policy.” A fluent final message should not erase those operational states.
Production does not mean fully autonomous
Autonomy is a risk and control choice, not a maturity ladder.
A production workflow may deliberately keep people responsible for:
- approving high-impact or irreversible actions;
- resolving ambiguous identity, policy, or evidence;
- reviewing low-confidence exceptions;
- responding to incidents and unknown side-effect status;
- changing permissions, routes, and acceptance thresholds.
Automation can expand when evidence supports it and shrink when risk or performance changes. A narrow system with well-enforced boundaries can be more production-ready than a broadly autonomous one with unclear authority.
A practical readiness review
Before broad release, a team should be able to answer:
- Which representative tasks and failure modes were evaluated?
- What user-visible indicators and targets define acceptable behavior?
- Which outputs and actions receive deterministic or independent validation?
- What can each agent and tool read, write, or execute?
- Which events make route, state, tool, validation, and cost failures observable?
- Which retries are safe, and how are uncertain side effects reconciled?
- What explicit states and limits stop the workflow?
- How does the system degrade, roll back, or escalate?
- Which decisions require human approval or review?
- Can an operator connect a user-visible failure to the responsible release, dependency, state transition, or control?
Passing this review does not guarantee perfect operation. It replaces an unsupported assumption – “the demo worked” – with inspectable evidence and a plan for what the system will do when reality differs from the script.
References
- Monitoring Distributed SystemsGoogle SRE
- Service Level ObjectivesGoogle SRE
- Elevated errors for Responses APIOpenAI Status, May 2026