Explainers
Mechanism-first explanations for readers who want to understand how a system behaves, not just what it is called.
- 01
The Round Trip: What Happens When You Hit Enter
A chat message becomes one part of an application-assembled request, a model generates from that request, and the application carries the result back.
- 02
The Load-Bearing Truths
Five mechanical facts replace the most misleading shortcuts about model generation, calls, context, memory, and the request a product actually sends.
- 03
Tokens Are Not Words
Text reaches a language model as learned vocabulary units, and those boundaries matter for context, cost, and system design.
- 04
Next-Token Prediction
A language model builds a response one token at a time, repeatedly scoring candidates against the context created so far.
- 05
Training and the Knowledge Cutoff
Training shapes a model's parameters before deployment; prompts and retrieved evidence shape a response at runtime without rewriting those parameters.
- 06
Why Models Hallucinate
A language model can produce a plausible claim without checking it against evidence; grounding and verification add that missing work but cannot guarantee truth.
- 07
Sampling and Temperature
Temperature reshapes the model's next-token distribution, while sampling turns those probabilities into choices that can send identical inputs down different paths.
- 08
Statelessness: No Memory Between Calls
A model call does not privately carry a conversation forward; continuity exists because an application, API, or framework supplies the relevant state again.
- 09
Anatomy of a Request and Response
A model call is a structured exchange: applications assemble instructions, content, controls, and tools, while responses separate generated output from protocol metadata.
- 10
The Real Payload: What Is Actually Transmitted
The text in the chat box is only one part of the effective request; instructions, history, tools, and injected context can all change behavior and consume budget.
- 11
The Context Window: The Central Constraint
Instructions, history, tools, user input, output, and sometimes reasoning all draw on finite per-call capacity, so adding context creates real tradeoffs.
- 12
Streaming: What You Are Actually Watching
Streaming changes when partial output or lifecycle events reach the application; it can improve perceived responsiveness without proving less total work or lower cost.
- 13
The Cost Model: Why You Pay What You Pay
The billable call can include instructions, replayed history, tools, generated output, and sometimes reasoning tokens, even when the latest user message is short.
- 14
From Answer to Action: Why One Call Isn't Enough
A model can propose a plan in one call, but controlled external work needs software that executes actions, returns observations, and decides whether to continue.
- 15
Tool Calling: How a Text Model "Does" Things
Tool use is a protocol boundary: the model proposes a structured call, external software decides whether to execute it, and a later model call receives the result.
- 16
The Agent Loop: Orchestration Around the Model
A practical agent is a controlled system around model calls; the orchestrator assembles state, dispatches tools, handles failures, and decides when to repeat or stop.
- 17
ReAct: Interleaving Reasoning and Action
ReAct names a reason-act-observe pattern in which external results can redirect the next decision, without requiring every system to expose reasoning or use the same control strategy.
- 18
Reading the Machine: What "Reviewed 8 Files" Really Means
Agent progress messages summarize selected model, host, and tool events; they are useful clues, not direct proof of comprehension, hidden reasoning, or completed verification.
- 19
The Memory Illusion: Memory as a Construction
AI products create continuity by storing, selecting, and reinjecting information around otherwise stateless model calls; the model itself does not own the persistent record.
- 20
How Meaning Becomes Math: Embeddings
Embeddings turn text into numeric vectors whose relative positions can rank related candidates; that ranking enables semantic search without proving that a result is true or sufficient.
- 21
RAG: Retrieval-Augmented Generation
RAG is an application-controlled pipeline that retrieves external evidence, places selected material into current context, and asks a model to generate from it without rewriting the model's parameters.
- 22
Why Retrieval Fails: Relevant ≠ Required
Similarity search can return material that is clearly about the topic yet omits the exact evidence an answer requires, creating a silent path from plausible context to a fluent mistake.
- 23
Persistent Knowledge vs. Query-Time Retrieval
Persistence describes where information remains available; retrieval describes how external information is selected for a request, so robust systems choose and combine knowledge paths by update, provenance, and failure needs.
- 24
Reasoning Models: What "Thinking" Tokens Actually Are
Reasoning-capable models can spend token-accounted work before or between visible outputs, while an API may expose only usage metadata or a summary; those artifacts reveal behavior without proving a complete causal trace or a correct answer.
- 25
The Faithfulness Problem: When the Explanation Is a Story
A model-generated explanation can be useful and an answer can be accurate without proving that the visible rationale faithfully describes the causal path that produced the result.
- 26
Confident and Wrong: Why Fluency Isn't Accuracy
Fluent wording and decisive explanations are properties of generated language, not evidence that a claim is correct; reliability must come from sources, recomputation, tests, live data, or review appropriate to the task.
- 27
Grounding and Verification: Engineering Trust
Trustworthy AI workflows connect claims to appropriate evidence, check that evidence outside the model's own explanation, and apply task-specific policies for acceptance, uncertainty, and escalation.
- 28
Why Long Conversations Degrade: Lost in the Middle
A long request can fit inside a model's context window and still fail because relevant evidence is buried, diluted by noise, stale, contradictory, or used unevenly.
- 29
Compaction and Summarization: Surviving the Window
Compaction keeps a long-running system within budget by preserving some information exactly, representing some more briefly, retrieving some later, and deliberately losing the rest.
- 30
Prompt Caching: Making Repetition Cheap
Prompt caching can reduce the cost and latency of processing a repeated request prefix, but it does not shrink context, select relevant information, create memory, or improve bad evidence.
- 31
The Rising Cost Curve: Why Sessions Get Expensive
In a long session, each later call may reconstruct more prior state than the one before it, so per-turn input, cumulative work, latency, and context pressure can rise even when new messages stay short.
- 32
Context Engineering: Managing the Budget Deliberately
Context engineering is the application-level discipline of selecting, representing, ordering, refreshing, and budgeting the information and capabilities a model receives for a particular task.
- 33
Multi-Agent Systems: Routing, Handoffs, Orchestration
Multi-agent systems compose familiar model and tool loops through explicit routing, delegation, ownership, and result contracts; extra agents help only when those boundaries provide a concrete benefit.
- 34
Subagent Isolation: Why the Orchestrator Is Half-Blind
Delegation gives a worker bounded context and a bounded return channel; the orchestrator can inspect only the results, artifacts, and telemetry the surrounding system makes available.
- 35
Demos vs. Production: The Reliability Gap
A demo proves that one trace can work; production requires the surrounding system to measure representative behavior, constrain actions, expose failures, recover safely, and operate under real variability.
- 36
The Trust Boundary: Prompt Injection and Security
Prompt injection is a system security problem: untrusted content can influence model behavior, so applications must keep permissions, validation, approval, and consequential execution outside the model's authority.
- 37
Mental Simulation: Thinking Like the Machine
Mental simulation turns an unfamiliar AI feature or failure into a testable system hypothesis by tracing payload, prediction, tools, state, context, orchestration, production controls, and trust boundaries.