Sampling and Temperature

Temperature reshapes the model's next-token distribution, while sampling turns those probabilities into choices that can send identical inputs down different paths.

  • Explainer
  • 5 min read
Illustration of temperature reshaping next-token probabilities before sampling selects an output.

Run the same prompt twice and you may receive different answers. That does not mean the model ignored the prompt or chose words without structure. It can be the expected result of sampling from a distribution of plausible next tokens.

Temperature changes the shape of that distribution. Sampling turns the reshaped probabilities into one selected token. Once selected, that token joins the context and changes every decision that follows.

From scores to a choice

At a generation step, the model produces a score, often called a logit, for each candidate token. A normalization function turns eligible scores into a probability distribution. A decoding policy must then decide which token becomes output.

Greedy decoding chooses the highest-scoring candidate. Sampling makes a weighted draw: high-probability candidates are more likely, but another eligible candidate can be selected.

Weighted does not mean arbitrary. If one token has much more probability mass than the others, it should appear more often over many comparable draws. It is still not guaranteed to appear on any single draw.

What temperature changes

Temperature is applied before sampling. In a common implementation, the candidate logits are divided by a positive temperature value before normalization.

Holding the logits and other decoding controls fixed:

  • a lower positive temperature magnifies score gaps, concentrating probability on the leading candidates;
  • a higher temperature compresses those gaps, distributing more probability to alternatives;
  • positive temperature scaling preserves the candidates’ score order, even though it changes how far apart their probabilities become.

Temperature does not add facts to the model, inspect whether a candidate is true, or improve the reasoning behind the scores. Its direct job is to reshape selection odds.

Same starting logits

Atlas 2.0 / Beacon 1.0 / Forge 0.0

Lower temperature / T = 0.5

More concentrated

  1. Atlas86.7%
  2. Beacon11.7%
  3. Forge1.6%

A likely draw: Atlas

Higher temperature / T = 2.0

More distributed

  1. Atlas50.6%
  2. Beacon30.7%
  3. Forge18.6%

A possible draw: Beacon

Selected token joins contextDifferent draw → different next distribution
Temperature reshapes selection odds.The probabilities come from a closed three-candidate teaching example. They are not measurements from a production model or probabilities that a candidate is true.

The example begins with three fictional codenames and fixed logits of 2.0, 1.0, and 0.0. At T = 0.5, most probability mass lands on Atlas. At T = 2.0, Atlas remains the leading candidate, but Beacon and Forge receive more meaningful chances.

These percentages are a transparent mathematical construction over three candidates. A production model scores a much larger vocabulary, and its decoding stack may also filter or alter candidates through top-k, top-p, repetition penalties, grammar constraints, or provider-specific rules.

Why one draw changes the rest

Suppose the lower-temperature run samples Atlas while the higher-temperature run samples Beacon. The difference is not confined to one word.

Each selected token is appended to the context:

Run A context: Name the deployment tool Atlas
Run B context: Name the deployment tool Beacon

The model now computes the next-token distribution from two different sequences. One branch may favor language about maps and navigation; the other may favor signals and detection. Later tokens diverge because the first different token changed the input used for every later prediction.

This is the same autoregressive feedback loop that builds any response. Sampling provides a branch point; append-and-repeat propagates its consequences.

Lower variation is not determinism

Lowering temperature often reduces run-to-run variation because less probability reaches lower-scored candidates. Observing the same answer several times is still not proof that the system is deterministic.

Exact or near-exact repetition depends on more than temperature:

  • the complete effective context, including hidden instructions and retrieved text;
  • the decoding strategy and every sampling parameter;
  • the random seed and random-number implementation, where exposed;
  • the model, tokenizer, and serving versions;
  • numerical kernels, hardware, batching, and other runtime behavior;
  • provider-side changes that may not be represented in the visible request.

A seed can control one source of randomness in a stable environment. It cannot promise identical output across arbitrary versions, platforms, or implementations. Numerical operations themselves may be nondeterministic, and a hosted API may not expose every relevant control.

If byte-for-byte identity is a requirement, treat it as a system property to test under pinned conditions, not a promise inferred from one parameter.

Temperature zero is not one universal contract

People often describe temperature = 0 as deterministic. The operational meaning depends on the API or library.

One implementation may reject zero and ask you to disable sampling for greedy decoding. Another may expose zero as a special value that selects the highest-scoring candidate. A provider can place additional behavior around either path.

The portable distinction is between sampling from a distribution and using a constrained selection rule such as greedy decoding. Read the target system’s current documentation before turning a conventional value into a guarantee.

Variation and correctness are separate

A stable answer can be consistently wrong. A varied set of answers can contain several valid ways to express the same supported conclusion.

Temperature therefore should not be treated as a truth dial:

  • lowering it does not verify a claim;
  • raising it does not give the model new knowledge;
  • variation does not by itself indicate hallucination;
  • repetition does not establish correctness.

For factual reliability, supply appropriate evidence and verify consequential claims. For output stability, constrain the decoding and execution environment. Those are related engineering goals, but they require different mechanisms.

Choosing controls from the requirement

Start with the behavior the application needs:

  • For reproducible tests, pin model and tokenizer versions, fix the full request, choose a constrained decoding policy, control randomness where possible, and allow for runtime limits documented by the platform.
  • For diverse ideation, permit sampling and evaluate the resulting candidates rather than assuming diversity implies quality.
  • For structured output, use grammar or schema constraints where available instead of expecting temperature to enforce syntax.
  • For factual claims, ground and verify the answer instead of attempting to tune truth through decoding.

The durable mental model is causal: temperature reshapes probabilities, sampling realizes one candidate, and that candidate changes the context for the next step. The rest of the response grows from the branch actually taken.

References

  1. The Curious Case of Neural Text DegenerationICLR, 2020
  2. Generation strategiesHugging Face Transformers
  3. GenerationConfigHugging Face Transformers
  4. ReproducibilityPyTorch