Multi-Head Latent Attention Explained
MLA compresses the KV cache into a small latent vector instead of sharing heads. Why it trades compute for memory, and which models bet on it.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
How large language models actually work, in plain terms.
MLA compresses the KV cache into a small latent vector instead of sharing heads. Why it trades compute for memory, and which models bet on it.
ReadA perfect needle-in-a-haystack chart says a model can find one planted sentence. It says very little about whether it can reason over your documents.
ReadPaged attention stores the KV cache in fixed-size blocks instead of one contiguous slab, which is what lets a server hold far more concurrent sessions.
ReadHow to size disk for open-weight models: what drives checkpoint size, why quantisations and versions multiply it, and why storage is rarely the real constraint.
ReadAttention is order-blind by default. How position gets injected, why the method decides how far context can stretch, and what you see when it fails.
ReadPretraining decides what a model knows. Post-training decides how it behaves. Knowing which stage owns a problem tells you whether prompting can fix it.
ReadFrequency and presence penalties suppress loops by punishing tokens the model already used. On code that punishes syntax, and the loop was rarely the real bug.
ReadThe one-line change that made hundred-layer networks possible, why the residual stream behaves like a shared bus, and what it explains about pruning.
ReadReinforcement learning from human feedback shapes the qualities nobody can write down. The three-stage pipeline, and the failure modes you see as a user.
ReadRoPE encodes position by rotating vectors rather than adding to them. Why that gives relative distance for free, and how it made 1M context possible.
ReadScaling laws describe how loss falls as compute, data and parameters grow. What they actually claim, where they stopped applying, and why it matters to you.
ReadSampling several answers and voting beats a single answer on some tasks and is pure waste on others. The mechanism, the cost, and when to reach for it.
ReadShowing 25–36 of 72 articles