Mixture of Experts: Why Trillion-Parameter Models Are Fast
MoE models decouple parameter count from compute per token, which is why a trillion-parameter model can be cheap to serve. Here is the mechanism and what it costs you.
Model announcements now routinely quote two parameter counts: a total and an active figure. GLM 5.2 is roughly 744B total with about 40B active per token. DeepSeek V4 Pro is 1.6T total with 49B active. Kimi K3 is 2.8T.
That gap is the entire point of mixture-of-experts architectures, and it explains why the largest models available are often not the slowest or most expensive to run.
The idea: decouple parameters from compute
In a dense transformer, every parameter participates in processing every token. Doubling the parameter count doubles the compute per token. That relationship is what made scaling expensive.
A mixture-of-experts layer breaks it. The feed-forward network in a transformer block is replaced by many parallel expert networks plus a small learned router. For each token, the router picks a handful of experts, and only those run. Everything else sits in memory unused for that token.
The Switch Transformer paper (Fedus, Zoph and Shazeer) made the extreme version of this work by routing each token to exactly one expert, which simplified routing enough to train trillion-parameter models with acceptable stability. Modern models generally route to several experts per token rather than one, but the principle is identical.
So total parameters describe how much knowledge the model can store, and active parameters describe how much work each token costs. These are now separate design dials.
How routing works
The router is a small learned layer that scores each expert for the current token and selects the top few. The token representation is passed through those experts, and their outputs are combined weighted by the router scores.
Nothing tells the experts what to specialise in. Specialisation emerges from training, and it is usually less interpretable than the marketing suggests — experts do not neatly divide into "the Python one" and "the French one". The routing patterns are real but messy.
The hard part of training MoE is load balancing. Left alone, the router collapses onto a few popular experts, leaving the rest undertrained and wasting the capacity you paid for. Training therefore adds an auxiliary loss that penalises imbalanced routing, and there is a genuine tension between routing tokens to the best expert and routing them to a less busy one.
What this means for serving
The consequences are asymmetric, and they explain a lot of the pricing you see.
Memory is set by total parameters. Every expert must be resident somewhere, because any token might route to it. A 744B-parameter model needs 744B parameters of weights in memory even though each token touches a fraction of them. This is why MoE models are hard to run on a single machine regardless of their active parameter count.
Compute is set by active parameters. Latency and throughput track the active figure. This is why models with enormous totals can be priced competitively — you are paying for the compute, not the storage.
Batching gets complicated. Different tokens in a batch route to different experts, so the work is irregular. Serving stacks handle this with expert parallelism, spreading experts across devices, which introduces communication between them. Getting good utilisation from an MoE deployment is materially harder than from a dense model.
Determinism suffers. Routing decisions can depend on how a batch is composed, which is one more reason identical requests do not always produce identical output on shared infrastructure.
What it means for you as a consumer
If you are calling an API, most of the above is somebody else's problem. Three things still matter.
- Do not read total parameters as capability. A 1.6T MoE and a 1.6T dense model would be enormously different objects. Compare on benchmarks and on your own evaluations, not on headline size.
- Do not read active parameters as capability either. The total is doing real work — the model genuinely has more capacity to store knowledge than its active count suggests.
- Cost per token is the number that matters. Architecture determines the vendor economics; price and quality on your task determine yours.
If you are self-hosting, the calculus flips. The memory footprint is the binding constraint, and an MoE model with a modest active count can still be out of reach on hardware that would happily run a dense model of the same speed. This is where quantisation earns its keep.
The trade-offs nobody advertises
MoE is not free capability. Several costs are real.
Parameter efficiency is worse. An MoE model with a given total parameter count is generally less capable than a dense model with the same total, because each parameter sees less of the training data. The bet is that the compute savings let you train a far larger total than you otherwise could, and that this wins overall. Empirically it does — but the comparison is not one-to-one.
Training is less stable. Routing collapse, load imbalance and gradient noise from discrete expert selection are ongoing engineering problems rather than solved ones.
Fine-tuning is trickier. Adapting an MoE model can disturb the routing distribution it learned, and the interaction between adapters and expert selection is less well understood than dense fine-tuning.
Quantisation behaves differently. Experts vary in sensitivity, so uniform low-bit quantisation across all of them is not always the right choice.
Where the field has landed
Sparse architectures now dominate the frontier for open-weight models, and the pattern is consistent: very large totals, modest active counts, long context. Kimi K3 (Moonshot AI, released 16 July 2026, open weights 27 July) sits at 2.8T parameters with a 1M-token context. GLM 5.2 (Z.ai, 13 June 2026) pairs roughly 744B total and 40B active with a 1M context and an MIT licence. DeepSeek V4 Pro reaches 1.6T total and 49B active, also MIT-licensed and notably cheap per token.
The practical upshot for a developer is unglamorous. Architecture is a reason the price-performance frontier moved, not a selection criterion in itself. Pick on measured quality for your task, on cost, and on latency — and treat the parameter counts as context for why those numbers look the way they do.
Common questions
What is the difference between total and active parameters?
Total is every parameter in the model and sets the memory needed to serve it. Active is the subset each token actually passes through, and it sets compute, latency and largely the price per token.
Are MoE models worse than dense models of the same size?
Per parameter, generally yes — each parameter sees less training signal. The trade pays off because the compute savings allow a far larger total than a dense model could reach at the same training budget.
Can I run an MoE model locally?
Only if you can hold all the weights in memory, since any token may route to any expert. The low active parameter count helps speed, not footprint, so quantisation is usually the deciding factor.