Why the Same Prompt Gives You a Different Answer
Temperature zero is not determinism, and a seed is only best effort. Here is where LLM nondeterminism really comes from and how to build around it.
You set temperature to 0, send the same prompt twice, and get two different answers. Nothing is broken. The assumption that greedy decoding implies reproducible output is simply wrong, and the reason is more interesting than "floating point is weird".
This matters beyond curiosity. Test suites that assert on exact strings, caches keyed on prompt hashes, and audit trails that claim a decision is reproducible all quietly depend on a guarantee no provider offers.
What temperature zero actually does
Temperature scales the logits before sampling. At zero, sampling collapses to argmax — always take the highest-probability token. That removes the sampling randomness, and only the sampling randomness.
It says nothing about whether the logits themselves are identical between two runs. If the numbers going into argmax differ in the last few bits, and two candidate tokens are close, argmax can flip. One flipped token changes the prefix for every token after it, and outputs diverge completely from there.
So temperature zero converts a stochastic process into one that amplifies tiny numerical differences instead of averaging over them.
Where the differences come from
The popular explanation is that floating-point addition is non-associative and GPU thread scheduling is unpredictable, so results wobble. Thinking Machines Lab examined this directly in their 2025 write-up Defeating Nondeterminism in LLM Inference and concluded that the usual story is mostly wrong: the individual kernels in a forward pass are typically run-to-run deterministic.
The real culprit is batching. Production servers group unrelated requests into batches whose size and composition depend on who else is calling at that moment. Kernels are not batch-invariant — the reduction order inside a matmul or a normalisation depends on batch shape — so the same request produces slightly different numerics depending on the traffic it happened to be batched with.
Their demonstration is worth remembering. Sampling 1,000 completions at temperature 0 from Qwen3-235B-A22B-Instruct-2507 with the same prompt produced 80 unique completions, the most common occurring 78 times. After making RMSNorm, matrix multiplication and attention batch-invariant, all 1,000 completions were identical — at a throughput cost, since their unoptimised deterministic path ran a benchmark in 55 seconds against vLLM's 26, improving to 42 with a better attention kernel.
Determinism is achievable. It is a deliberate engineering trade against throughput, and hosted endpoints have not generally made that trade.
The other sources, in rough order of impact
- Model updates. A version alias silently repointed at new weights is the largest single cause of "it changed overnight".
- Serving stack changes. Kernel upgrades, different GPU generations in a heterogeneous fleet, changed quantisation.
- Mixture-of-experts routing. Expert selection can be affected by what else is in the batch, adding another batch-dependent path.
- Sampling parameters. Any temperature above zero, plus top_p and top_k, are randomness by design.
- Your own prompt. Injected timestamps, unordered retrieval results and map iteration order make prompts differ when you believe they are identical.
That last one is worth checking before you blame the model. Diff two request bodies byte for byte; the number of times the answer is on your side is high.
What seed and system_fingerprint buy you
OpenAI exposes a seed parameter described as making a best effort to sample deterministically, alongside a system_fingerprint identifying the backend configuration that served the request. Same seed, same parameters and same fingerprint should mostly give the same output.
Read the hedging honestly. Determinism is not guaranteed, variability is still observed in practice, and the fingerprint changes whenever the provider updates the numerical configuration behind the model. Seed narrows the distribution; it does not collapse it.
The fingerprint is still useful — as a change detector. Log it with every request and you get an early warning that the backend moved, which is exactly the information you want when quality shifts for no apparent reason.
Build for variance instead of fighting it
The productive move is to stop treating variance as a defect and start treating it as a property with a measurable size.
- Never assert on exact strings. Assert on properties: valid JSON, required fields present, the correct tool selected, a numeric answer within tolerance.
- Use structured outputs where the shape matters. Constrained decoding removes format variance even when wording varies.
- Run every eval case k times. A single run tells you nothing about consistency. Five to ten tells you a lot.
- Report the strict metric. If a workflow needs all steps to succeed, measure the fraction of cases where all k attempts succeeded, not the fraction where at least one did. A 90 percent per-attempt success rate is only about 57 percent consistent across eight attempts.
- Pin the model version. Call the dated snapshot, not the moving alias, and upgrade deliberately.
The only true determinism is a cache
If you genuinely need identical output for identical input — a legal disclosure, a regulated calculation, a golden-file test — do not ask the model. Store the output.
Generate once, review it, and serve the stored artefact thereafter. Regenerate only on a deliberate version bump. This is the standard approach for anything where reproducibility is a requirement rather than a convenience, and it sidesteps the entire problem.
When determinism is worth the cost
Two cases justify pursuing it properly. Reinforcement learning, where a mismatch between the sampling policy and the training policy turns an on-policy algorithm into an off-policy one without anyone noticing. And debugging, where being able to reproduce a failure exactly is the difference between fixing it and guessing.
Both require control over the serving stack, which in practice means self-hosting with batch-invariant kernels. For application work, the better investment is an evaluation suite that measures variance rather than an architecture that pretends it does not exist.
Common questions
Does temperature 0 guarantee the same output?
No. It removes sampling randomness but not numerical variation. When two candidate tokens are close, small differences from batching can flip the argmax and the output diverges from that point.
What is the seed parameter actually for?
It narrows variability on a best-effort basis. Combined with a stable system fingerprint, repeated requests mostly match, but providers explicitly do not guarantee identical output.
How do I write tests against a nondeterministic model?
Assert on properties rather than exact text, run each case several times, and report how often all attempts succeeded. For output that must never change, cache a reviewed response and serve that.