Temperature, Top-p and Sampling: What the Knobs Really Do
Temperature is not a creativity dial and top-p is not a quality setting. Here is what each sampling parameter changes mathematically, and how to set them for real work.
Every generation step ends the same way: the model produces a score for every token in its vocabulary, and something has to pick one. That picking step is sampling, and it is the only part of inference you control directly at request time.
Most of the folklore around it is wrong. Temperature is not a creativity slider, top-p is not a quality setting, and setting temperature to zero does not guarantee you the same output twice.
Logits, softmax, and where temperature enters
The model outputs raw scores called logits — one per vocabulary entry, unbounded and not yet probabilities. Softmax exponentiates them and normalises so they sum to one.
Temperature divides the logits before that softmax. That single division is the whole mechanism:
- Temperature below 1 divides by a small number, magnifying gaps between logits. The distribution sharpens and the top token dominates.
- Temperature of 1 leaves the distribution exactly as the model produced it.
- Temperature above 1 flattens the distribution, lifting the tail and making unlikely tokens more reachable.
Note what it does not do. It never adds an option the model had not already assigned mass to. High temperature does not generate new ideas — it makes the model more willing to take the second- or fifth-best continuation, which sometimes reads as originality and sometimes reads as incoherence.
Top-p is truncation, not scaling
Nucleus sampling, introduced by Holtzman et al. at ICLR 2019 in The Curious Case of Neural Text Degeneration, attacks a different problem. Their observation was that maximisation-based decoding such as beam search produces bland, repetitive text, while unrestricted sampling occasionally draws from a very long tail of low-probability garbage.
Top-p sorts tokens by probability, walks down the list accumulating mass, and cuts off once the cumulative total reaches p. Sampling then happens only within that nucleus, renormalised.
The important property is that the cutoff is dynamic. When the model is confident, the nucleus might contain one or two tokens. When it is genuinely uncertain, it might contain hundreds. The paper notes the nucleus typically ranges from one to around a thousand candidates depending on context.
Top-k is the fixed-size cousin: keep the k highest-probability tokens regardless of how the mass is distributed. It is cruder, because k=40 is far too permissive when the model is certain and far too restrictive when it is not.
Min-p and why it exists
Top-p degrades at high temperature. Once you have flattened the distribution enough, the cumulative mass threshold admits a lot of noise.
Min-p, introduced in Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs (July 2024), sets the threshold relative to the top token instead. If min-p is 0.05 and the leading token has probability 0.6, anything below 0.03 is dropped. The filter tightens automatically when the model is confident and relaxes when it is not.
The paper reported improvements in both quality and diversity across Mistral and Llama 3 models at higher temperatures, with human raters preferring its output. It has since been adopted in Hugging Face Transformers, vLLM and other serving stacks, so it is often available even when the vendor API does not expose it.
Do not stack them without thinking
Temperature and top-p interact, and turning both up is how people end up with output that is simultaneously erratic and repetitive.
A workable discipline: pick one primary knob. If you want more variety, raise temperature and leave top-p at a permissive default. If you want to suppress tail noise, tighten top-p and leave temperature at 1. Changing both at once means you cannot attribute the resulting behaviour to either.
Repetition and frequency penalties are a separate axis. They subtract from the logits of tokens that have already appeared. They are blunt instruments — they penalise legitimately repeated tokens too, which in code means variable names and keywords. Use them sparingly outside of free-form prose.
Temperature zero is not determinism
Setting temperature to 0 means greedy decoding: always take the highest-probability token. That removes the sampler as a source of variation, and it is genuinely the right default for extraction, classification and structured output.
It does not make the API deterministic. Floating-point reductions on GPUs are not associative, so results depend on batch composition and kernel scheduling, both of which vary with server load. Mixture-of-experts routing can add further batch-dependent variation. Two identical requests can therefore diverge, usually after a long shared prefix, because one token flipped and everything after it conditioned on the flip.
If you need reproducibility, the tools are a fixed seed where the vendor offers one, plus pinning the exact model version. Treat identical output as a strong tendency, not a contract, and never build a cache key or a test assertion that assumes byte equality.
Settings that actually work
- Structured output, extraction, classification: temperature 0. You want the mode of the distribution, and diversity is pure downside.
- Code generation: low, typically 0 to 0.3. There are usually few correct answers and many plausible-looking wrong ones.
- Agent loops and tool selection: low. A creative choice of tool is a bug, and errors compound across turns.
- Summarisation and rewriting: moderate, around 0.3 to 0.7. Some variation in phrasing is fine; variation in facts is not.
- Brainstorming and drafting: higher, and prefer generating several samples over pushing a single sample to an extreme setting.
Two caveats worth internalising. First, reasoning models often ignore or constrain these parameters, because the sampling regime is part of how the model was trained to think; check the vendor documentation rather than assuming your defaults apply. Second, if output quality is bad at temperature 0, sampling is not your problem — the prompt or the model choice is.
How to tune without guessing
Build a small evaluation set of real inputs with known-good outputs. Sweep one parameter across three or four values, run each several times, and score the results. You are looking for two things at once: average quality and variance.
The common finding is that the useful range is narrower than expected, and that most perceived gains from sampling tweaks disappear once you measure across more than a handful of examples. That is a useful result — it redirects effort to prompt structure and model selection, where the leverage actually is.
Common questions
Should I change temperature or top-p?
Change one. Temperature reshapes the whole distribution; top-p truncates its tail. Adjusting both at once makes the effect impossible to attribute, and stacking high values on each other produces erratic output.
Does temperature 0 give identical output every time?
No. It removes sampling randomness, but GPU floating-point reductions are not associative and results vary with batch composition and server load. Use a fixed seed and a pinned model version, and still do not assert byte equality.
What is min-p and when should I use it?
A dynamic threshold set relative to the top token probability rather than cumulative mass. It holds up better than top-p at high temperature, so it is worth trying for creative generation on stacks like vLLM that expose it.