Small vs Large Models: Why Parameter Count Stopped Meaning Much
Models

Small vs Large Models: Why Parameter Count Stopped Meaning Much

Mixture-of-experts split model size into total and active parameters, and only one of them predicts your bill. How to think about size when picking a model.

"How big is it" used to be one question. With mixture-of-experts architectures it is two, and they predict different things.

Kimi K3 is 2.8 trillion parameters and activates roughly 104 billion of them per token. DeepSeek V4 Pro is 1.6 trillion and activates 49 billion. MiniMax M3 is 428 billion and activates about 23 billion. The totals span a factor of six and a half; the active counts span a factor of four and a half, in a different order of magnitude entirely.

Total parameters tell you what it costs to hold the model in memory. Active parameters tell you what it costs to generate a token. If you are buying inference from an API, the second number is the one connected to your bill and your latency.

What each number actually predicts

Total parameters predict serving feasibility. Kimi K3's weights are about 1.56TB, which means a multi-node deployment with fast interconnect before you serve a single request. That is why "open weights" for the largest models means provider competition rather than self-hosting.

Active parameters predict throughput and price. The relationship is visible in measurement. Artificial Analysis clocks MiniMax M3, at roughly 23B active, generating around 112 tokens per second with a 1.5-second time to first token. DeepSeek V4 Pro, at 49B active, runs around 74 tokens per second at 1.8 seconds. Roughly double the active parameters, roughly two thirds the throughput.

Neither reliably predicts capability. DeepSeek V4-Flash is 284B total with 13B active — an order of magnitude smaller than V4 Pro on both counts — and lands within a few points of it on practical coding benchmarks, at around $0.14 and $0.28 per million tokens against $0.435 and $0.87. A July 2026 re-post-training pushed Flash to 82.7 on Terminal-Bench 2.1 with no price change. Training data and post-training method now move capability more than size does.

Where small models genuinely win

The Qwen3.5 line is the clearest evidence, because it spans real sizes under Apache 2.0: a 27B dense model, 35B-A3B, 122B-A10B and a 397B-A17B flagship. Reported figures in the smaller range land in the same band as small closed models — around 72 on SWE-bench Verified for the 27B, near 70 for 35B-A3B, and 72.2 on BFCL-V4 tool use for 122B-A10B against 55.5 for GPT-5 mini. Sourcing varies on the exact digits, but the shape is consistent.

Small models win on four things that large models cannot buy back:

  • Latency. In interactive tools, time to first token determines whether the thing feels responsive. No amount of capability fixes a two-second pause on every keystroke-adjacent completion.
  • Iteration count. Faster loops mean more attempts, and more attempts frequently beat one better answer — particularly in agent workflows with a test to check against.
  • Deployability. A 27B model fits on hardware you can own. That is the only route to weights that never leave your network.
  • Cost at volume. Bulk classification, extraction, commit messages and test scaffolding do not need frontier reasoning, and running them on a frontier model is pure waste.

Where large models still earn it

The gap is not uniform across task types, and this is the part that matters for routing.

Large models pull ahead on ambiguity, on long horizons, and on the tail of genuinely hard problems. Kimi K3 sits around 57 on the Artificial Analysis Intelligence Index against the mid-forties for DeepSeek V4 Pro and MiniMax M3, and ranked first on Arena Frontend Code at 1,679 in blind developer voting. GLM-5.2, at around 40B active, reports 81.0 on Terminal-Bench 2.1 — the first open-weight model past 80 on long agentic sessions when it launched in June 2026, up from 62.0 a generation earlier.

Note that GLM-5.2 achieves that with fewer active parameters than DeepSeek V4 Pro. Size is not what produced the Terminal-Bench result; targeted post-training on long-horizon tasks is. That is the strongest available argument against choosing by parameter count.

The honest summary is that a small model degrades gracefully on easy work and fails sharply on hard work, and the boundary between those is codebase-specific. You cannot read it off a leaderboard.

Finding your own boundary

The measurement is straightforward and takes an afternoon:

  1. Take twenty tasks from your git history, spread across bug fixes, small features, refactors and explanations.
  2. Run all twenty on the smallest credible model. Score pass, salvageable or fail, and record turns to completion.
  3. Re-run only the failures on the next size up. Then re-run whatever still fails on the largest.
  4. Record the escalation rate at each tier. That is your routing policy, expressed as a number.

Most teams find the small model clears the large majority. Knowing exactly which minority it does not clear is worth more than any aggregate score, because it converts model choice from an argument into a lookup.

Two things that break the intuition

Reasoning effort is a second size dial. GLM-5.2 exposes high and xhigh; DeepSeek V4 exposes high and extended. Turning effort up spends more output tokens and adds latency, and it changes results more than a size step sometimes does. Kimi K3 launched with a single effort level, which removes that lever entirely.

Caching flattens the input side. Kimi K3 charges $0.30 per million on cached input against $3.00 on a miss. If you send a large stable prefix every call, the cost gap between a large and a small model narrows sharply on input and remains wide on output — which means the right optimisation is often to shorten responses, not to shrink the model.

Common questions

Is a model with more total parameters always better?

No. GLM-5.2 reports 81.0 on Terminal-Bench 2.1 with about 40B active parameters, ahead of models with more. Post-training method and data now move capability more than size, and total parameters mainly predict serving cost.

What is the difference between total and active parameters?

In a mixture-of-experts model, total parameters are what must be held in memory; active parameters are what actually run for each token. Active count drives throughput, latency and price. Total count drives whether you can serve it at all.

How do I know when a small model is not enough?

Measure it. Run twenty tasks from your git history on the small model, re-run only the failures on a larger one, and record the escalation rate. That number is your routing policy and it is specific to your codebase.

Similar articles

Best Model for Low Latency: Time to First Token Wins
Models
Models·8 min read

Best Model for Low Latency: Time to First Token Wins

For interactive work, perceived speed is decided by time to first token, not throughput. Which models and settings actually make an interface feel fast.

Read
Latency-Adjusted Model Scoring: When Fast Beats Smart
Models
Models·9 min read

Latency-Adjusted Model Scoring: When Fast Beats Smart

Benchmarks measure what a model answers, never how long it took. How to weight latency into model selection for interactive and agentic workloads.

Read
A/B Testing Two Models Without Fooling Yourself
Models
Models·9 min read

A/B Testing Two Models Without Fooling Yourself

Comparing two models on live traffic sounds simple and usually is not. Sample sizes, paired designs, and the metrics that actually settle the question.

Read