Speculative Decoding: Faster Output, Identical Tokens
Fundamentals

Speculative Decoding: Faster Output, Identical Tokens

Speculative decoding drafts several tokens cheaply and verifies them in one pass, cutting latency without changing the output distribution. How it works.

Generating text one token at a time is slow for a reason that has nothing to do with arithmetic. At low batch sizes, decoding a single token requires streaming the entire model weights from memory to compute, and that memory traffic dominates. The GPU spends most of its time waiting, not calculating.

Speculative decoding exploits that idle capacity. Guess several tokens cheaply, then check all of them in one forward pass of the real model, at close to the cost of checking one. When the guesses are right you got several tokens for the price of a step. When they are wrong you fall back and lose almost nothing.

The mechanism

Two models. A small fast draft model and the large target model you actually want output from.

  1. The draft model generates a short continuation — typically a handful of tokens — autoregressively. It is small, so this is cheap.
  2. The target model scores all of those draft tokens in a single forward pass, in parallel, because the whole candidate sequence is already known.
  3. A verification rule walks the draft left to right and accepts tokens while they are consistent with what the target would have produced. At the first rejection, it discards the remainder and samples a corrected token from an adjusted distribution.
  4. Repeat.

The step that makes this interesting is number three. The acceptance rule is a modified rejection sampling scheme, and it is constructed so that the tokens coming out are distributed exactly as if the target model had generated them alone.

The output does not change

This is the property that separates speculative decoding from every other inference optimisation, and it is worth being precise about.

Quantisation changes the output. Distillation changes the output. Pruning changes the output. Speculative decoding does not: the modified rejection sampling preserves the target model distribution, within hardware numerics. It is a pure latency optimisation.

Practically, that means you do not need to re-run your evaluations after enabling it. Sample-level output can differ run to run exactly as much as it already did with temperature above zero, and no more.

Where it came from

Two papers landed within months of each other and established the technique.

Leviathan, Kalman and Matias published Fast Inference from Transformers via Speculative Decoding (arXiv 2211.17192, ICML 2023), reporting 2x to 3x acceleration on T5-XXL against the standard T5X implementation, with identical outputs and no retraining or architecture change.

Chen and colleagues at DeepMind published Accelerating Large Language Model Decoding with Speculative Sampling (arXiv 2302.01318) in February 2023, benchmarking on the 70B-parameter Chinchilla and reporting a 2x to 2.5x decoding speedup in a distributed setup without compromising sample quality or modifying the model.

Both make the same core observation: scoring a short continuation in parallel costs about as much as sampling one token, so cheap guesses are nearly free to check.

Acceptance rate is the whole economics

The speedup is governed by how many draft tokens survive verification on average — the acceptance length.

That number depends on how well the draft model imitates the target on this text. It is not constant. Boilerplate, repeated identifiers, closing brackets, formulaic prose and the mechanical parts of code all draft extremely well. Genuinely novel content drafts poorly. The original insight is exactly this: hard language modelling tasks contain easy subtasks that a small model approximates well.

Draft length is a tuning parameter with a real optimum. Draft too few tokens and you leave speedup on the table. Draft too many and you spend draft compute on tokens that get rejected. Most systems settle in the low single digits, adaptively where the runtime supports it.

Variants that avoid a separate draft model

Running two models is operationally awkward — two sets of weights, two things to keep in memory, and a draft model that has to share the target vocabulary. Several approaches remove that requirement.

Extra prediction heads. Medusa-style designs attach additional heads to the target model that predict several positions ahead, then verify the resulting candidate tree. No second model to serve.

Feature-level drafting. The EAGLE line drafts in the target model feature space rather than in token space. EAGLE-3 draws on low, mid and high level features from the target, which improves how often draft tokens are accepted; reported acceptance lengths average in the region of 2.8 tokens per verification step across a broad domain mix.

N-gram and lookup drafting. No neural draft model at all. Propose continuations by matching against text already in the context. This works remarkably well when output repeats input — editing a file, restating a schema, refactoring code — and costs essentially nothing.

The batch size caveat

This is the part most write-ups skip, and it decides whether speculative decoding helps you.

The technique converts spare compute into fewer sequential steps. That trade is excellent when decoding is memory-bandwidth-bound, which is the single-request, low-batch regime. As batch size rises, the GPU becomes compute-bound, the spare capacity disappears, and verifying rejected tokens starts costing real throughput.

The published numbers show the pattern clearly. One study of draft-head methods reported roughly 2.70x throughput improvement at batch size one, falling to 1.63x at batch size eight. EAGLE-3 measurements show around 2.3x at batch size four, declining towards break-even by batch size thirty-two.

So: speculative decoding is a latency optimisation for interactive workloads, not a throughput optimisation for bulk ones. On a saturated batch server it can make things worse.

What this means if you are just calling an API

You do not control this directly, and you should not need to. It is worth understanding for three reasons.

First, it explains why streaming output speed varies within a single response. Predictable stretches stream in bursts; novel stretches slow down. That is acceptance rate changing, not the network.

Second, it explains why a provider can improve latency without changing the model. Enabling speculative decoding is invisible in output quality by construction.

Third, it is a useful mental model for agent latency. Agent loops are mostly low-concurrency, latency-sensitive workloads with highly predictable output — tool call syntax, repeated file paths, structured JSON. That is close to the best case for speculation, which is part of why per-step latency in agentic tools is often better than raw model size would suggest.

Common questions

Does speculative decoding reduce output quality?

No. The verification step uses a modified rejection sampling rule that preserves the target model distribution within hardware numerics, so the tokens produced are distributed exactly as if the large model had generated them alone.

How much faster is speculative decoding?

The founding papers reported 2x to 3x on T5-XXL and 2x to 2.5x on a 70B Chinchilla model at low batch size. The gain shrinks as batch size grows, because decoding stops being memory-bandwidth-bound and the spare compute the technique relies on disappears.

Why does it need a draft model at all?

It needs cheap candidate tokens from somewhere. Newer variants avoid a separate model by adding extra prediction heads to the target, drafting in the target feature space, or matching n-grams already present in the context.

Similar articles

Speculative Decoding in Practice: When It Pays Off
Fundamentals
Fundamentals·9 min read

Speculative Decoding in Practice: When It Pays Off

Speculative decoding is standard in serving stacks now, but the speedup is workload-dependent. What decides whether it helps you, and how to tell.

Read
LLM Latency: Why Time to First Token Is the Number to Watch
Fundamentals
Fundamentals·9 min read

LLM Latency: Why Time to First Token Is the Number to Watch

LLM latency is two different problems wearing one name. Understanding prefill and decode tells you which knob to turn when a request feels slow.

Read
Throughput vs Latency: The Trade-off Behind Your Bill
Fundamentals
Fundamentals·9 min read

Throughput vs Latency: The Trade-off Behind Your Bill

Serving more tokens per second and serving them faster are opposing goals. The knob that reconciles them is batch size, and it decides both speed and price.

Read