Tokens Are Not Words

Text reaches a language model as learned vocabulary units, and those boundaries matter for context, cost, and system design.

  • Explainer
  • 4 min read
Illustration of text being broken into token fragments rather than whole words.

Before a language model can continue a sentence, the text has to become something the model can process. That conversion is tokenization.

It is tempting to treat a token as a word, or perhaps as a fixed handful of characters. Both shortcuts are useful until they are not. The actual unit comes from a tokenizer-specific vocabulary learned from patterns in training text. A boundary can land inside a word, include a leading space, join punctuation to nearby text, or split an identifier into several pieces.

That is not a cosmetic implementation detail. Tokens consume the context window, determine how much text fits in a request, and often form the basis of model pricing.

A vocabulary learned from patterns

Many language-model tokenizers use a subword method related to byte-pair encoding. The rough construction is mechanical:

  1. Start with small representable units.
  2. Count which adjacent units occur together frequently in a corpus.
  3. Merge frequent pairs into reusable vocabulary entries.
  4. Repeat, producing a vocabulary that represents common patterns compactly while leaving less familiar patterns fragmented.

The result is not a dictionary. It has no obligation to respect grammar, syllables, or the spaces a person sees between words. It reflects which sequences were useful to represent in the data used to build that tokenizer.

This is why a familiar suffix might be one unit in one context while an unusual identifier is broken into many. Byte-level approaches also ensure arbitrary input remains representable, including unfamiliar scripts and symbols, but their boundaries can look especially unintuitive around whitespace and punctuation.

Why the usual estimate fails

You may have seen a planning heuristic such as “one token is about four English characters.” For ordinary English prose and rough estimates, that can be directionally useful. It is not a conversion rule.

Compare these two requests:

Open the configuration file and set the timeout.
Open ./config/app.prod.json; set requestTimeout_ms=1200.

The second line contains path separators, punctuation, mixed casing, an underscore, digits, and a unit. Human eyes still recognize a short instruction. A tokenizer sees statistical patterns from its vocabulary, not “one instruction with roughly the same meaning.”

Code, hashes, numeric strings, markup, repeated punctuation, and mixed scripts are all places where a prose-calibrated estimate can drift. Even a leading space can affect which vocabulary entry matches. If a request limit, chunk size, or cost calculation matters, inspect the complete payload with the tokenizer used by the target model family.

Equal meaning does not imply equal cost

Two sentences can communicate the same idea while producing different token sequences. Languages and scripts have different coverage in a learned vocabulary. Structured content and prose have different recurring patterns. Neither visible length nor semantic equivalence fixes the result.

This does not mean one language has an inherent, timeless token cost. Tokenizer vocabularies change. Model families use different tokenizers. A comparison that is accurate for one version can become stale for the next.

The durable engineering rule is simpler: measure representative input rather than promoting a heuristic into a limit.

Where the unit shows up

Once text becomes tokens, the same unit appears throughout the system:

  • Context: instructions, conversation history, tool definitions, retrieved text, and generated output must fit within the model and runtime’s enforced limits; exact input and output accounting varies.
  • Cost: providers commonly meter some combination of input, output, cached, and reasoning tokens.
  • Chunking: retrieval pipelines split source material partly to control how many tokens enter a request.
  • Latency: longer token sequences create more input to process, while generated output grows one token at a time.
  • Reliability: a limit based on characters can pass local tests and still fail on production payloads with different languages or structure.

Common misconception

“A token is basically a word.”

A token is one entry from a tokenizer’s vocabulary. Some entries resemble words. Others are word fragments, punctuation patterns, spaces, bytes, or combinations that have no clean linguistic name. Treating all of them as words hides the mechanism precisely where engineering constraints begin to matter.

A useful prediction

Suppose two support prompts carry the same request. One is short English prose. The other mixes Hebrew, a product identifier, a file path, and a numeric error code.

Without running the relevant tokenizer, neither prompt can be declared cheaper from appearance alone. It is reasonable to expect the mixed and structured payload to fragment more, but that remains a tokenizer-specific prediction. The decision becomes evidence only after measurement.

That distinction is the point: tokens are not a visual property of text. They are the output of a particular representation process.

References

  1. Neural Machine Translation of Rare Words with Subword UnitsAssociation for Computational Linguistics
  2. Language Models are Unsupervised Multitask LearnersOpenAI
  3. Tokenizers documentationHugging Face
  4. tiktokenOpenAI