How LLMs Actually Work: Next-Token Prediction, End to End
Fundamentals

How LLMs Actually Work: Next-Token Prediction, End to End

A large language model does one thing repeatedly. Here is what happens between your prompt and the first token out, and what that mechanism explains about behaviour.

A large language model does exactly one thing: given a sequence of tokens, it produces a probability distribution over which token comes next. Chat, code generation, tool calling and agent loops are all that single operation run repeatedly, with the output appended to the input each time.

Holding that in mind explains most of the behaviour developers find surprising — why the model is confident when it is wrong, why phrasing changes results, and why the same request costs different amounts on different days.

The architecture in one pass

Nearly every model in production descends from the transformer, introduced in Attention Is All You Need (Vaswani et al., NeurIPS 2017). The paper replaced recurrence with attention entirely, which is what made training parallelisable enough to scale.

Your text goes through four stages:

  1. Tokenisation. The string is split into subword units and each unit is mapped to an integer ID.
  2. Embedding. Each ID becomes a vector. Position information is added, because attention itself is order-blind.
  3. Stacked transformer blocks. Each block runs self-attention followed by a feed-forward network, with residual connections around both.
  4. Output projection. The final vector is projected back to vocabulary size and softmaxed into probabilities.

Then a sampler picks one token from that distribution, the token is appended, and the whole thing runs again.

What attention is doing

Self-attention lets every token look at every other token in the sequence and decide how much each one matters to it. Concretely, each token emits a query, a key and a value; the query is compared against all keys to produce weights, and the output is the weighted sum of values.

The practical consequence is that meaning is contextual rather than fixed. The token for set in a sentence about tennis and set in a sentence about Python data structures start from the same embedding and end up in very different places by the final layer.

Multi-head attention runs several of these in parallel with different learned projections, so one head can track syntax while another tracks long-range coreference. Nobody assigns those roles; they emerge from training.

Where the knowledge lives

Attention moves information around. The feed-forward layers are where most of the parameters sit, and they behave more like a large associative lookup — patterns learned during training get retrieved and mixed in.

This matters because it explains a hard limit: a model knows what was compressed into its weights during training, and nothing else. It has no access to your database, no awareness of today, and no mechanism for checking whether a recalled pattern is true. Retrieval and tool calling exist precisely to route around that.

Modern frontier models increasingly make those feed-forward layers sparse. A mixture-of-experts model routes each token to a small subset of expert networks, so total parameter count and per-token compute decouple. GLM 5.2 (Z.ai, released 13 June 2026) is roughly 744B total parameters with about 40B active per token; DeepSeek V4 Pro is 1.6T total with 49B active. The memory footprint is the total; the speed is closer to the active figure.

Pretraining and post-training are different animals

Pretraining is the expensive part: predict the next token across an enormous text corpus. What comes out is a model that can continue text plausibly but will happily continue a question with three more questions, because that is what documents do.

Post-training is what makes it usable. Supervised fine-tuning on instruction-response pairs teaches the model to answer rather than continue. Preference optimisation — reinforcement learning from human or model feedback — then shapes tone, refusal behaviour and formatting.

Two things follow. First, the personality differences between models are largely post-training, not architecture. Second, most of what you experience as a model being helpful or annoying is a policy decision someone made, not a property of transformers.

Prefill and decode: why latency behaves oddly

Inference splits into two phases with completely different performance profiles.

Prefill processes your entire prompt at once. It is highly parallel and compute-bound, and it determines your time to first token. Doubling the prompt roughly doubles this work, and more than doubles it as attention costs grow.

Decode generates tokens one at a time. Each step needs the full attention state for every previous token, which is cached — the KV cache — so it is memory-bandwidth-bound rather than compute-bound. This is why output tokens are usually priced higher than input tokens: they are generated sequentially and cannot be batched within a single request.

It also explains prompt caching. If a long prefix is identical between requests, the KV cache for that prefix can be reused, skipping most of prefill. That is a real saving, and it requires the prefix to be byte-identical, so keeping timestamps and session IDs out of your system prompt is worth actual money.

What the mechanism explains

  • Confident errors. The model outputs a distribution over tokens, not a truth value. A fluent wrong answer and a fluent right answer are produced by the same process.
  • Prompt sensitivity. Different wording puts you in a different region of the distribution. This is not the model being fussy; it is the model being conditional.
  • No arithmetic guarantees. Digits are tokens. Multi-digit multiplication is pattern completion unless the model is given a calculator.
  • Recency effects. Everything is conditioned on the sequence so far, and the most recent tokens have the strongest influence on the next one.

The practical takeaway

Treat the model as a conditional next-token engine and your design choices get simpler. Put the important material close to the question. Give it tools for anything requiring ground truth or exact computation. Keep prompt prefixes stable so caching works. And when output quality is inconsistent, look at sampling settings and prompt structure before concluding the model cannot do the task.

Common questions

Do LLMs understand what they are saying?

They build rich contextual representations that support genuinely useful behaviour, but the training objective is next-token prediction, not truth. Fluency and correctness are produced by the same mechanism, which is why verification has to come from outside the model.

Why are output tokens more expensive than input tokens?

Input is processed in one parallel prefill pass. Output is generated one token at a time, each step bound by memory bandwidth rather than compute, so it is far harder to amortise across a batch.

What is the KV cache?

The stored attention keys and values for tokens already processed. It lets each new token attend to the history without recomputing it, and it is what makes prompt caching across requests possible when the prefix is identical.

Similar articles

Grouped-Query Attention: Why Long Context Got Affordable
Fundamentals
Fundamentals·9 min read

Grouped-Query Attention: Why Long Context Got Affordable

GQA shrinks the KV cache by sharing keys and values across query heads. What it costs in quality, and why it is on the spec sheet of nearly every model.

Read
Attention Mechanisms Explained Without the Linear Algebra
Fundamentals
Fundamentals·9 min read

Attention Mechanisms Explained Without the Linear Algebra

What attention actually computes, why it made transformers work, and why its cost scaling explains almost every practical limit you hit with long context.

Read
Beam Search vs Sampling: Why Chat Models Do Not Search
Fundamentals
Fundamentals·9 min read

Beam Search vs Sampling: Why Chat Models Do Not Search

Beam search finds higher-probability text and worse text. Why sampling won for open-ended generation, and where search-like decoding still earns its place.

Read