Tokenization Explained: Why Your Bill Is Not About Words
Fundamentals

Tokenization Explained: Why Your Bill Is Not About Words

Tokenizers decide what your prompt costs, how long your context really is, and why models miscount letters. Here is how subword tokenization works in practice.

Models do not read characters and they do not read words. They read tokens — subword fragments produced by a tokenizer that was fitted to a training corpus before the model ever existed.

That layer is invisible until it is not. It sets your cost, it decides how much actually fits in a context window, and it is behind a whole category of bugs that look like reasoning failures but are not.

Byte-pair encoding, briefly

Most production tokenizers use a variant of byte-pair encoding. The fitting procedure is simple: start with individual bytes, count adjacent pairs across a corpus, merge the most frequent pair into a new symbol, and repeat until you hit a target vocabulary size.

What comes out is a vocabulary where frequent sequences are single tokens and rare ones get split. Common English words are usually one token. An unusual surname might be four. Whitespace is typically attached to the following word, which is why the and the are different tokens.

Byte-level BPE matters here: because the base alphabet is bytes rather than characters, the tokenizer can represent anything at all, including emoji and text in scripts it never saw. It just represents unfamiliar text badly, at several tokens per character.

Vocabulary size is a trade-off, not an upgrade

OpenAI shipped cl100k_base with GPT-4 and GPT-3.5 Turbo, at roughly 100k tokens, then roughly doubled it to about 200k with o200k_base for the GPT-4o generation. Larger vocabularies compress text into fewer tokens, particularly for non-English scripts.

The cost is on the other side. A bigger vocabulary means a bigger embedding table and a bigger output projection, both of which scale with vocabulary size. It also means rarer tokens get fewer training updates each. Vendors pick a point on that curve; there is no universally correct answer, and different model families land in different places.

A practical consequence: token counts are not portable. The same document can differ by a double-digit percentage between two vendors, so a cost estimate built against one tokenizer is only an estimate against another.

Not all text costs the same

English prose is the best case, at roughly four characters per token. Everything else is worse, and some of it is much worse.

  • Code fragments badly. Indentation, punctuation runs and camelCase identifiers all split. Budget closer to three characters per token.
  • JSON is expensive relative to its information content. Braces, quotes and repeated key names are pure overhead paid on every record.
  • Base64 and hashes are close to worst case — near-random character sequences have no frequent pairs to merge.
  • Non-English text pays what researchers call a tokenization tax. Benchmarking work on cross-lingual tokenization has measured length inflation of up to roughly 15x across languages for the same content, driven largely by multibyte scripts and underrepresentation in the fitting corpus.

That last point has a direct commercial edge. When billing is per token and latency scales with token count, users writing in some languages pay materially more for the same request. Recent work on parity-aware BPE explicitly optimises against that disparity, but the models you are using today were fitted before it.

The bugs tokenization causes

A cluster of well-known model failures are tokenizer artefacts rather than reasoning failures.

Counting letters. Asking how many times a letter appears in a word is hard because the model never sees letters. It sees two or three opaque chunks. The information is recoverable, but it is not directly available in the way the question implies.

Reversing strings and character-level edits. Same cause. Anything that operates below the token boundary is working against the representation.

Arithmetic. Numbers tokenize inconsistently depending on digit grouping, so the same quantity can have different internal representations in different contexts. This is part of why long multiplication is unreliable without a tool.

Trailing whitespace. Ending a prompt with a space can put the model in an awkward spot, because it now has to produce a continuation token that would normally have carried its own leading space. Strip trailing whitespace from prompts.

The fix for all of these is the same: if the task is genuinely character-level or arithmetic, give the model a tool. Do not prompt harder at a representational limitation.

Counting tokens before you spend them

Rough estimates are fine for planning and useless for enforcement. Two habits are worth building.

First, count with the actual tokenizer for the model you are calling. Every major SDK exposes one, and several vendors expose a token-counting endpoint. Do it before the request, not by reading the usage field afterwards.

Second, log input token counts per request in development. Almost every team that does this discovers a prompt several times larger than they assumed, usually because a tool schema, a retrieved document set, or a serialised config object is being pasted into the system prompt on every call.

Then set a ceiling in your own code and fail loudly when a prompt exceeds it. A runaway context in an agent loop compounds, because every subsequent turn carries the bloat forward.

Design decisions that follow

  1. Choose compact formats for bulk data. Sending tabular data as CSV rather than an array of JSON objects removes the repeated key names, which is often a third of the payload.
  2. Shorten identifiers in generated payloads where they are machine-consumed. Not in your source code — in the blobs you send to and from the model.
  3. Truncate tool output aggressively. The first 50 matches instead of all 1,200 changes the cost of the entire remaining session.
  4. Do not paste hashes, UUID lists or base64 into context. They are maximally expensive and almost never load-bearing.
  5. Benchmark cost in your own language. If your users write in Thai or Japanese, per-token pricing means something different to you than it does to the vendor benchmark.

None of this requires understanding merge tables. It requires accepting that the unit you are billed in is not the unit you think in, and measuring the gap once rather than guessing at it repeatedly.

Common questions

How many tokens is a page of text?

Roughly 600 to 800 tokens for a dense technical page in English. Code is denser per character, so a 200-line source file typically lands around 2,000 to 3,000 tokens.

Why can a model not count the letters in a word?

It never sees the letters. The word arrives as one or two subword tokens with no explicit character breakdown, so character-level questions ask about information the representation does not directly expose.

Are token counts the same across different models?

No. Vocabulary and merge rules differ per tokenizer, so the same text can differ by a substantial margin between vendors. Always count with the tokenizer belonging to the model you are actually calling.

Similar articles

Byte-Pair Encoding Explained: How Tokenizers Are Built
Fundamentals
Fundamentals·9 min read

Byte-Pair Encoding Explained: How Tokenizers Are Built

BPE is a compression algorithm that became the standard way to split text for language models. How it is trained, what it produces, and why it behaves oddly.

Read
Token Costs by Language: Why the Same Text Bills Differently
Fundamentals
Fundamentals·9 min read

Token Costs by Language: Why the Same Text Bills Differently

The same sentence costs more in some languages than in English, and code has its own profile. Where the multiplier comes from and how to measure yours.

Read
Tokenizer Differences: Why the Same Text Costs Different Amounts
Fundamentals
Fundamentals·8 min read

Tokenizer Differences: Why the Same Text Costs Different Amounts

Two models quoting the same price per million tokens can cost different amounts for identical text. Why tokenizers vary and what it does to your bill.

Read