Quantization Explained: Big Models on Smaller Hardware
Quantization shrinks model weights to fewer bits so they fit in less memory. Here is how the formats differ, what accuracy you lose, and how to pick one without guessing.
Model weights are numbers, and numbers can be stored at different precisions. Quantization is the practice of storing them in fewer bits — usually 8 or 4 instead of 16 — so the model occupies less memory and moves less data per token.
The reason this matters is that inference is largely memory-bandwidth-bound during generation. Halving the bytes you read per token does not just make a model fit; it makes it faster.
The mechanics, minus the maths
A weight stored in 16-bit floating point takes two bytes. Stored as a 4-bit integer it takes half a byte, a 4x reduction. To recover an approximate original value you keep a scale factor, and often a zero point, for each small group of weights.
Group size is the hidden dial. Per-tensor scaling is cheapest and least accurate; per-group scaling, typically over 32 to 128 weights, costs a little metadata and recovers most of the fidelity. When two 4-bit quantizations of the same model differ noticeably in quality, group size and calibration are usually why.
The rough footprint arithmetic is worth memorising. Multiply parameter count by bytes per parameter: FP16 is 2, INT8 is 1, 4-bit is about 0.5 plus overhead. Then add space for the KV cache, which grows with context length and batch size and is frequently the thing that actually blows your memory budget.
Weights, activations and the KV cache
These are three separate targets and people conflate them.
Weight-only quantization is the common case. Weights are stored in low precision and dequantized on the fly for computation. It cuts memory and bandwidth without needing hardware support for low-precision arithmetic.
Weight and activation quantization goes further, running the actual matrix multiplications in low precision. This requires hardware that supports the format natively, and it is where the large throughput gains come from. NVIDIA NVFP4 is the current example — a 4-bit floating-point format with native acceleration on Blackwell tensor cores that quantizes both weights and activations, reported to approach FP8 accuracy with substantially higher throughput than weight-only 4-bit methods.
KV cache quantization is separate again and often overlooked. In long-context or high-batch serving the cache can rival the weights for memory. Quantizing it to 8 bits is usually low-risk; going to 4 bits tends to degrade long-context recall, which is exactly what you wanted the long context for.
The formats you will actually encounter
- GGUF — the llama.cpp format, used by Ollama and most desktop tooling. Mixed-precision integer quantization with a family of variants. Q4_K_M is the widely recommended default. Its distinguishing feature is CPU and Apple Silicon support, and the ability to offload some layers to GPU and keep the rest in system RAM.
- GPTQ — calibration-based post-training quantization to 4-bit integers, using sample data to minimise reconstruction error layer by layer. GPU-oriented, well supported in serving stacks.
- AWQ — activation-aware weight quantization. Same family as GPTQ, but it identifies which weights matter most by examining activations and protects those. It tends to edge out GPTQ on reasoning-heavy evaluations.
- FP8 — 8-bit floating point with native support on recent data-centre GPUs. Near-baseline quality at meaningfully higher throughput, which makes it the default for production serving where the hardware supports it.
- bitsandbytes — the on-the-fly option, and the one behind QLoRA. Convenient because it needs no separate quantization step, generally slower at inference than a purpose-built kernel.
Independent comparisons converge on a consistent picture: 4-bit post-training quantization with AWQ or GGUF Q4_K_M is the practical sweet spot for most deployments, with code-generation degradation in the low single-digit percentages against the unquantized baseline. Optimised 4-bit kernels can also out-throughput FP16, because the bandwidth saving outweighs the dequantization cost.
What you actually lose
The honest summary is that the loss is small on average and unevenly distributed.
8-bit quantization is close to free for most purposes. 4-bit is a real but usually acceptable step down. Below 4 bits, degradation accelerates sharply and the model starts failing in ways that are hard to characterise.
More importantly, average benchmark scores hide where the damage lands. Quantization tends to hurt hardest on long chains of reasoning, where small errors compound; on rarely-seen knowledge, which was encoded in weights with little redundancy; on exact formatting and structured output; and on non-English text. A model that loses two points on a general benchmark can lose considerably more on your specific task.
The rule that follows: a larger model quantized to 4 bits usually beats a smaller model at full precision, when both fit. But you must verify on your own workload rather than trusting the aggregate.
Choosing without guessing
- Work out your memory ceiling first. Parameters times bytes per parameter, plus KV cache for your intended context length and concurrency. This eliminates most options immediately.
- Prefer a bigger model at 4-bit over a smaller one at 16-bit if both fit and speed is acceptable.
- Match the format to the runtime. GGUF for llama.cpp, Ollama, CPU and Apple Silicon. AWQ or GPTQ for vLLM and similar GPU servers. FP8 or NVFP4 if the hardware supports it natively.
- Do not go below 4 bits unless you have measured the specific task and accepted the result.
- Quantize the KV cache to 8 bits before touching the weights further if long context is what is squeezing you.
- Evaluate on your own inputs. Twenty representative examples run against both the quantized and unquantized model tell you more than any published table.
When not to bother
If you are calling a hosted API, quantization is the provider's decision and not visible to you. It is one reason the same model name can behave slightly differently across providers, which is an argument for evaluating providers as well as models — but it is not a knob you turn.
Quantization is a tool for people running weights themselves: local development, on-premise deployment, cost control at scale, or fitting a model onto hardware you already own. If none of those describe you, the interesting question is which model and which provider, not how many bits.
Common questions
How much quality do I lose at 4-bit?
Typically a few percent on general benchmarks with a good method like AWQ or GGUF Q4_K_M. The loss concentrates on long reasoning chains, rare knowledge and non-English text, so measure on your own task rather than trusting the average.
Is a smaller unquantized model better than a bigger quantized one?
Usually not. A larger model at 4-bit generally outperforms a smaller model at full precision when both fit in memory and latency is acceptable. Verify on your workload, since the gap narrows on tasks that quantization hits hardest.
Which quantization format should I use?
Match it to your runtime. GGUF for llama.cpp, Ollama, CPU or Apple Silicon; AWQ or GPTQ for GPU serving stacks like vLLM; FP8 or NVFP4 where the hardware accelerates them natively.