Open-Weight Models in 2026: Licences, Sizes and Trade-offs
Kimi K3, GLM-5.2, DeepSeek V4, MiniMax M3 and Qwen compared on the axes that decide deployment: licence terms, active parameters, serving cost and benchmark shape.
"Open weight" is treated as a binary in most comparison tables. It is not. Across the 2026 field the licences range from Apache 2.0 with a patent grant to custom documents with revenue-triggered obligations, and the serving requirements range from a single GPU to a multi-node cluster.
Those two axes decide whether a model is usable in your situation far more often than a benchmark score does. Here is the field arranged along them.
The licence spectrum
From most permissive to least:
- Apache 2.0 — the Qwen open-weight releases, including the Qwen3.5 family from February 2026. Permissive redistribution plus an explicit patent grant, which is the clause corporate legal teams actually ask about.
- MIT — GLM-5.2 and DeepSeek V4 Pro. No revenue thresholds, no regional restrictions, no separate agreements. Fine-tune, redistribute and resell freely.
- Custom community licence — MiniMax M3. Weights are published on Hugging Face, but under bespoke terms rather than an OSI-approved licence. Read them before shipping.
- Custom licence with revenue gates — Kimi K3. Attribution required above roughly 100 million monthly active users or $20 million monthly revenue, and a separate agreement with Moonshot required for any Model-as-a-Service business exceeding $20 million in aggregate revenue over a twelve-month period. Moonshot deliberately says "open weight", never "open source".
If you are an internal engineering team, all four tiers behave identically. If your product is inference, the bottom tier introduces a conversation with a vendor at precisely the moment you succeed, and the top two do not.
Serving footprint: total versus active parameters
Total parameters set your memory floor. Active parameters set your compute per token. Mixture-of-experts designs have pulled these far apart, and quoting only the total is how comparison tables mislead.
- Kimi K3 — 2.8T total, roughly 104B active. The weights are about 1.56TB. Multi-node for anyone.
- DeepSeek V4 Pro — 1.6T total, 49B active. An activation ratio near 3%.
- GLM-5.2 — around 744B total (Z.ai cites 753B), about 40B active.
- MiniMax M3 — 428B total, roughly 23B active. The leanest forward pass of the group.
- DeepSeek V4-Flash — 284B total, 13B active.
- Qwen3.5 family — 397B-A17B at the top, down through 122B-A10B and 35B-A3B to a 27B dense model.
Only the bottom of that list is realistically self-hostable by a normal team. For everything above it, the practical value of open weights is provider competition and price pressure, not local deployment — which is still worth a great deal, but is a different thing from what people usually mean.
Capability, and the shape of it
On the Artificial Analysis Intelligence Index the open field currently spreads from around 44 to around 57: Kimi K3 highest at roughly 57, GLM-5.2 around 51, and DeepSeek V4 Pro and MiniMax M3 clustered in the mid-forties. For reference, the top closed models sit near 60.
The index compresses four weighted categories into one number, so equal scores do not mean equal models. The benchmark each lab chose to publish tells you more:
- Kimi K3 — first on Arena Frontend Code at 1,679 in blind developer voting, ahead of Claude Fable 5; 1547 Elo on the Artificial Analysis long-horizon knowledge-work evaluation.
- GLM-5.2 — 81.0 on Terminal-Bench 2.1 (82.7 best reported), 62.1 SWE-bench Pro, 74.4 FrontierSWE, 99.2 AIME 2026. Long unattended agent runs.
- DeepSeek V4 Pro — 80.6% SWE-bench Verified, 93.5 LiveCodeBench, 3206 Codeforces, 90.1 GPQA Diamond. Verifiable problems.
- MiniMax M3 — 59.0 SWE-bench Pro, 66.0 Terminal-Bench 2.1, 74.2 MCP-Atlas, plus native image and video input.
Most of those are vendor-run on vendor scaffolds, which reliably score above standardised harnesses. The Arena result is the exception and is correspondingly harder to game.
Price, which is where open weights have actually won
The competitive effect is visible in the rate cards. DeepSeek V4 Pro is around $0.435 per million input tokens and $0.87 per million output after a permanent 75% cut in May 2026, with V4-Flash near $0.14 and $0.28. MiniMax M3 lists around $0.24 and $0.96 under discounting, tiered upward above 512K input tokens. GLM-5.2 is around $1.40 and $4.40 officially, with third-party providers lower. Kimi K3 is $3 and $15, with cached input at $0.30.
That is a spread of more than fifty times per output token across models that are all described as frontier-adjacent. No closed vendor faces that kind of pressure on its own weights, and it is the single clearest consequence of open releases.
How to actually choose
Work through these in order and the field narrows fast:
- Licence. Will you resell inference or redistribute a fine-tune? If so, eliminate anything with a revenue gate before evaluating capability.
- Deployment. Must the weights run inside your network? If so, only the Qwen sizes and V4-Flash are realistic, and the conversation is over.
- Task shape. Verifiable problems, long unattended agent runs, interactive code a human will judge, or visual inputs. Each has a different leader.
- Cost per completed task. Not per token. Measure turns and total spend on your own eval set.
Then keep switching cheap. All of these are served over OpenAI-compatible endpoints, so keep model identifiers in configuration rather than code and keep an eval set current. The field moved substantially between April and July 2026, and it will move again before the year ends.
Common questions
Which open-weight model has the most permissive licence?
The Qwen open-weight releases under Apache 2.0, which adds an explicit patent grant. GLM-5.2 and DeepSeek V4 Pro under MIT are effectively equivalent for most purposes. Kimi K3 is the most restricted, with revenue-triggered obligations.
Does open weight mean I can run it myself?
Rarely, at the top of the range. Kimi K3 is about 1.56TB of weights and needs a multi-node cluster. In practice open weights mean multiple providers can serve the model competitively, which is what has driven prices down so sharply.
How close are open models to closed frontier models?
Closer than a year ago and not level. On the Artificial Analysis Intelligence Index the strongest open model sits around 57 against roughly 60 for the leading closed models, and Kimi K3 ranked first on Arena Frontend Code in blind voting.