GPU Memory for Inference: What Actually Fills the Card
Weights are only the first line of the memory budget. Where the rest goes, why concurrency runs out before compute does, and how to size a deployment.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
How large language models actually work, in plain terms.
Weights are only the first line of the memory budget. Where the rest goes, why concurrency runs out before compute does, and how to size a deployment.
ReadGQA shrinks the KV cache by sharing keys and values across query heads. What it costs in quality, and why it is on the spec sheet of nearly every model.
ReadTraining is a one-off capital cost you never pay. Inference is a recurring cost you pay per request. How the two differ and why the distinction shapes pricing.
ReadA pretrained model continues text; it does not answer questions. Instruction tuning is the small, cheap stage that turns one into the other.
ReadThe KV cache is why generation is fast and why long context is expensive in memory rather than compute. What it stores and what it costs you.
ReadWhy deep transformers need normalisation, what pre-norm changed, and why RMSNorm won. The connection to quantisation and low-precision serving.
ReadAn MoE model only delivers its promised capacity if traffic spreads across experts. Collapse during training and skew during serving are both real risks.
ReadA model does not emit words, it emits a score for every token in its vocabulary. What softmax does to those scores and why probabilities are not confidence.
ReadAccuracy does not hold flat and then fall off a cliff at the context limit. It declines gradually, and the decline starts far earlier than most teams assume.
ReadA response that ends mid-word is almost never a model failure. It is a cap you set, or a default you never set. How truncation actually happens.
ReadTensor, pipeline and expert parallelism split a model across devices in different ways. Each moves a different cost onto the network, and that decides latency.
ReadParameter counts once tracked capability closely and no longer do. What size still predicts, what it never predicted, and what to check instead.
ReadShowing 13–24 of 72 articles