What Is a Parameter Count, and Does It Matter?
Fundamentals

What Is a Parameter Count, and Does It Matter?

Parameter counts are the most quoted and least understood model spec. What they measure, why total and active differ, and when the number predicts anything useful.

A parameter is a single learned number inside the model — one weight in one matrix. Training adjusts billions of them until the network produces useful continuations of text. The parameter count is simply how many of those numbers exist.

It gets quoted like a clock speed, and it has the same problem clock speeds had: it measures one input to performance, not performance. Two models with identical parameter counts can differ enormously in capability, and a model with a quarter of the parameters can beat one four times its size.

What the number physically means

Parameters are what you store and what you multiply. That gives the count two direct, non-negotiable consequences.

Memory. Every parameter occupies space determined by its numeric precision:

bytes = parameters × bytes_per_parameter

FP16 / BF16   2 bytes    per parameter
FP8           1 byte     per parameter
INT4          0.5 bytes  per parameter

So a 70B model at 16-bit precision needs roughly 140 GB just for weights, before any working memory. At 4-bit it is closer to 35 GB. This is arithmetic, not a benchmark, and it is the single most useful thing the parameter count tells you.

Compute per token. Generating one token requires roughly two floating-point operations per active parameter. More active parameters means more compute, which means higher latency and higher cost to serve. Providers price accordingly, which is why parameter count correlates with price even when it correlates poorly with quality.

Total parameters versus active parameters

This is the distinction that makes most modern parameter counts confusing, and you cannot read a spec sheet without it.

A dense model uses every parameter for every token. A mixture-of-experts model splits much of the network into many expert subnetworks and routes each token through only a few of them. The model stores everything but computes with a fraction.

Two published examples make the gap concrete. GLM 5.2, released by Z.ai in June 2026, is a mixture-of-experts model with roughly 744B total parameters and about 40B active per token. DeepSeek V4 Pro is roughly 1.6T total with about 49B active. In both cases you must hold the full weights in memory, but you pay compute closer to the active figure.

The practical reading:

  • Total parameters tell you what it costs to host — memory, and therefore how many GPUs.
  • Active parameters tell you what it costs to run — compute per token, and therefore latency and price.
  • A headline number without the split is close to uninterpretable for a MoE model. Kimi K3, at 2.8T parameters with open weights released in July 2026, is a very different hosting proposition from a 2.8T dense model that does not exist.

Why the count predicts capability badly

Capability depends on at least four things the parameter count says nothing about.

Training data. Quantity, quality, deduplication and domain mix. A model trained on more and better tokens beats a larger model trained on less.

Post-training. Instruction tuning, reinforcement learning from feedback, and reasoning training change behaviour far more than a size bump does. Two checkpoints of the same base model can differ by a wide margin on the same benchmark.

Architecture. Attention variants, routing strategy, context handling. These determine how efficiently the parameters are used.

Distillation. Small models trained on the outputs of large ones inherit a meaningful share of the teacher behaviour at a fraction of the size. This is now a standard technique and it decouples size from quality further every year.

The empirical evidence is straightforward: composite capability rankings routinely place models with very different parameter counts adjacent to each other, and closed models do not publish counts at all yet clearly sit at the top of those rankings. If the number were predictive, that would not happen.

When the count is genuinely useful

It is not a useless spec. It is a useful spec for a narrow set of questions:

  • Can I run this locally? Parameters times bytes-per-parameter, plus KV cache, versus your available memory. This is the number one reason to care.
  • How many GPUs to self-host? Same arithmetic, then divide by per-card memory and add overhead.
  • Roughly what should this cost per token? Active parameters set a floor on serving cost. A provider offering a very large active-parameter model far below the market rate is doing something you should ask about — quantisation, batching trade-offs, or subsidy.
  • Is quantisation likely to hurt? Larger models generally tolerate aggressive quantisation better than small ones, which degrade noticeably below about 4 bits.

What to look at instead

If your question is "will this model do my job well", parameter count is not the input. Use, in order:

  1. Your own evaluation set. Twenty real tasks from your workload, scored consistently. This beats every published number.
  2. Composite indices that aggregate many benchmarks, which are noisy but less gameable than any single score.
  3. Task-specific benchmarks in your domain, read with the awareness that contamination is common and leaderboards are optimised against.
  4. Serving characteristics — context window actually available, latency, throughput, tool-calling reliability.

The mental model worth keeping: parameter count is a capacity ceiling, not an achievement. It tells you how much the model could in principle have learned and exactly what it costs to store. What it actually learned is a question only evaluation answers.

Common questions

Does a higher parameter count mean a better model?

Not reliably. Training data, post-training and architecture matter more, and distillation lets small models inherit much of a large model behaviour. Parameter count sets a capacity ceiling and a hosting cost, not a capability level.

What is the difference between total and active parameters?

In a mixture-of-experts model, total parameters are what you must store in memory, while active parameters are the subset used to generate each token. Total drives hosting cost, active drives compute, latency and price per token.

How much memory does a model need for its parameters?

Multiply parameters by bytes per parameter: 2 bytes at 16-bit precision, 1 byte at FP8, half a byte at 4-bit. A 70B model is therefore roughly 140 GB at 16-bit or 35 GB at 4-bit, before the KV cache and runtime overhead.

Similar articles

Model Size vs Capability: Why Bigger Stopped Meaning Better
Fundamentals
Fundamentals·9 min read

Model Size vs Capability: Why Bigger Stopped Meaning Better

Parameter counts once tracked capability closely and no longer do. What size still predicts, what it never predicted, and what to check instead.

Read
Mixture of Experts: Why Trillion-Parameter Models Are Fast
Fundamentals
Fundamentals·8 min read

Mixture of Experts: Why Trillion-Parameter Models Are Fast

MoE models decouple parameter count from compute per token, which is why a trillion-parameter model can be cheap to serve. Here is the mechanism and what it costs you.

Read
Model Distillation Explained
Fundamentals
Fundamentals·8 min read

Model Distillation Explained

How a small model inherits the behaviour of a much larger one, what gets lost along the way, and why distillation is the reason cheap models got good so fast.

Read