LLM Latency: Why Time to First Token Is the Number to Watch
Fundamentals

LLM Latency: Why Time to First Token Is the Number to Watch

LLM latency is two different problems wearing one name. Understanding prefill and decode tells you which knob to turn when a request feels slow.

"The model is slow" is never one problem. It is either a long wait before anything appears, or a steady trickle once it does. Those two symptoms have different causes and different fixes, and conflating them is why latency work so often goes nowhere.

The split follows directly from how inference runs. A request is processed in two phases with completely different performance characteristics.

Prefill and decode

Prefill reads your entire prompt and builds the KV cache. Every input token can be processed in parallel, so this phase saturates the GPU and is compute-bound. It ends when the first output token appears, which makes prefill the thing that sets time to first token.

Decode then generates output one token at a time. Each step depends on the previous one, so there is no parallelism within a request — and each step re-reads the model weights and a growing KV cache from memory. Decode is memory-bandwidth-bound, and it sets the pace of everything after the first token.

Two phases, two bottlenecks. That single fact explains most of what follows.

The four numbers worth measuring

  • TTFT — time to first token. Perceived responsiveness. Dominated by prompt length and queueing.
  • TPOT — time per output token, also called inter-token latency. How fast the stream flows once it starts.
  • Output tokens per second — the reciprocal of TPOT, which is how most dashboards report it.
  • End-to-end latency — roughly TTFT plus TPOT multiplied by the remaining output tokens.

That last formula is the useful one, because it shows where your time actually goes. A short answer is a TTFT problem. A long answer is a TPOT problem. A long answer to a long prompt is both, and you should fix them separately.

What makes TTFT bad

Prompt length is the first cause, and it is close to linear. The same model answering a 500-token question and a 100,000-token document analysis has very different first-token waits, because prefill has to ingest everything before it can emit anything.

Queueing is the second. Under load your request waits for a batch slot before prefill even begins. This is why TTFT percentiles look fine in a quiet test and terrible in production — the median measures the model, the tail measures the queue.

Prompt caching is the strongest available fix. When a prefix has recently been processed, the provider can reuse that work instead of recomputing it. OpenAI states that prompt caching can cut time-to-first-token latency by up to 80 percent on qualifying prompts. It only helps if your prefix is genuinely stable, which is a design decision, not a setting.

What makes TPOT bad

Decode speed is mostly a property of the model and the serving stack, not of your request. Larger models move more weight bytes per token. Mixture-of-experts architectures help here, because only a fraction of the parameters are active per token — that is a large part of why very large MoE models can serve at reasonable speeds at all.

The counter-intuitive part is batching. Providers batch many concurrent requests to use the hardware efficiently, which raises total throughput while slightly slowing any individual stream. Throughput and per-request latency trade against each other, and you are usually buying somebody else's throughput optimisation.

The lever you actually control is output length. Decode cost is linear in tokens emitted. Asking for a concise answer, or a structured one with a bounded shape, is often the single most effective latency change available to an application developer.

Reasoning tokens are decode time

Models that produce extended internal reasoning before answering are spending decode steps you cannot see. From the harness perspective this shows up as a long gap before user-visible content, even though the model is generating the whole time.

If a reasoning model feels slow, the question is not "why is inference slow" but "how many thinking tokens is this task actually worth". Where the API exposes an effort or budget control, that is the correct dial. Where it does not, task decomposition is the fallback.

Agents multiply everything

A chat turn pays latency once. An agent pays it per iteration, and iterations are serial by construction: the model cannot request the next tool until it has seen the last result.

Worse, each iteration prefills a longer prompt than the one before, because history accumulates. Ten turns into a session your TTFT can be several times what it was at turn one, for the same model and the same hardware.

Three things help materially. Execute independent tool calls concurrently rather than in sequence — most models can request several per turn, and many harnesses still run them one at a time. Keep the prompt prefix stable so caching survives across iterations. And truncate tool output aggressively, since every token you keep is re-prefilled on every subsequent call.

Streaming changes the experience, not the work

Streaming does not make a response faster. It moves the moment the user sees progress from end-to-end latency to TTFT, which is usually the difference between "slow" and "fine".

It also changes what you should optimise. With streaming enabled, TTFT and TPOT are your user-facing metrics and total time barely matters. Without it, only end-to-end latency exists. Decide which mode you are in before you tune anything.

Measure it properly

  • Measure at your client, not at the provider. Network, gateway and your own middleware are part of what users experience.
  • Report p50, p95 and p99 separately for TTFT and TPOT. Averages hide the queueing that causes complaints.
  • Bucket by prompt length. A single TTFT number across mixed traffic is meaningless when prefill scales with input size.
  • Log cache hit rate alongside TTFT. Unexplained TTFT regressions are very often a prefix that stopped matching.
  • Track tokens per request, not just requests. Latency regressions frequently turn out to be prompt growth.

Then fix in this order: cut input tokens, stabilise the cache prefix, cut output tokens, parallelise tool calls, and only after all of that consider a smaller or faster model.

Common questions

Does a bigger context window make requests slower?

A large window costs nothing on its own. Filling it does. Prefill scales with the tokens you actually send, so time to first token grows with prompt length regardless of the advertised limit.

Why is the first token slow but the rest fast?

They are different phases. Prefill processes the whole prompt before anything is emitted and is compute-bound; decode then emits one token at a time and is limited by memory bandwidth.

What is the cheapest latency win for an agent?

Running independent tool calls concurrently and truncating tool output. Both reduce serial round trips and stop history growth from inflating prefill on every later iteration.

Similar articles

Time to First Token: What It Measures and What Moves It
Fundamentals
Fundamentals·9 min read

Time to First Token: What It Measures and What Moves It

TTFT is dominated by prefill compute, queueing and network distance rather than by model speed. What each contributes, and how to measure it without fooling yourself.

Read
Speculative Decoding: Faster Output, Identical Tokens
Fundamentals
Fundamentals·8 min read

Speculative Decoding: Faster Output, Identical Tokens

Speculative decoding drafts several tokens cheaply and verifies them in one pass, cutting latency without changing the output distribution. How it works.

Read
Speculative Decoding in Practice: When It Pays Off
Fundamentals
Fundamentals·9 min read

Speculative Decoding in Practice: When It Pays Off

Speculative decoding is standard in serving stacks now, but the speedup is workload-dependent. What decides whether it helps you, and how to tell.

Read