Kimi K3: What a 2.8T Open-Weight Model Actually Buys You
Models

Kimi K3: What a 2.8T Open-Weight Model Actually Buys You

Moonshot shipped the largest open-weight model yet: 2.8T parameters, 1M context, and a licence that is not MIT. A technical read on where K3 earns its price.

Moonshot AI released Kimi K3 on 16 July 2026 and published the weights ten days later, on 26 July. At 2.8 trillion total parameters it is the largest open-weight model anyone has shipped — the download is roughly 1.56TB.

Size is the headline, and it is also the least interesting thing about the model. What matters for anyone deciding whether to route work to it is the shape of its capability, the licence attached to the weights, and whether the price per task justifies itself. Those three things point in different directions.

What is actually in the model

K3 is a mixture-of-experts model. The 2.8T figure is the total parameter count; reporting on the open-weight release puts active parameters at roughly 104B per token, so the compute per forward pass is far smaller than the headline suggests. That ratio — large total, moderate active — is the standard trade in 2026: you pay for the memory to hold the model, not for dense compute on every token.

The architecture uses Kimi Delta Attention plus attention residuals, and Moonshot describes a stable latent-MoE routing scheme with MXFP4/MXFP8 quantisation baked into the release. The context window is 1M tokens, and vision is native rather than bolted on.

One practical quirk at launch: K3 shipped with a single reasoning effort level, described as "max". Models like GLM-5.2 and DeepSeek V4 expose two or more effort settings so you can trade latency for depth. With K3 you get one setting, which removes a lever you may be used to pulling.

Where the evidence says it is strong

The most credible signal is not a self-reported table. Arena ranked K3 first on its Frontend Code evaluation at 1,679 points, ahead of Claude Fable 5, in blind head-to-head developer voting. Blind preference on real prompts is a much harder number to game than a benchmark you score yourself.

Artificial Analysis reported that on their private long-horizon knowledge-work evaluation, K3 reached an overall Elo of 1547 — an improvement of 732 points over Kimi K2.6, and behind only Claude Fable 5. They also noted K3 used 21% fewer output tokens than K2.6 to get there, with a cost per task of $0.94 against $1.04 for GPT-5.6 Sol. On the Artificial Analysis Intelligence Index, K3 sits around 57, which is the highest of any open-weight model.

Moonshot's own numbers put K3 ahead of Claude Opus 4.8 max and GPT-5.5 high on most tasks, and behind Claude Fable 5 and GPT-5.6 Sol. They cite 91.2 on BrowseComp and 1668 Elo on GDPval-AA v2. Treat vendor scaffolds with the usual scepticism — a lab tuning its own harness routinely beats the same model under a standard one — but the direction is corroborated by the independent evals above.

Moonshot itself flagged a UX gap against Fable 5 and GPT-5.6 Sol despite the benchmark numbers. That is an unusually honest admission and worth taking seriously: raw capability and pleasant behaviour in a long session are different properties.

The licence is the part to read

K3 is open weights, not open source, and Moonshot is careful to use that phrasing. The weights ship under a custom "kimi-k3" licence rather than the Modified MIT terms that came with K2.

Two clauses matter. Products or services above roughly 100 million monthly active users or $20 million monthly revenue carry an attribution requirement. More significantly, any Model-as-a-Service business with aggregate revenue over $20 million across any twelve-month period needs a separate agreement with Moonshot to use the model commercially.

For an internal engineering team this is irrelevant. If you are building an inference product on top of these weights, it is the first thing to send to legal — and it is a genuine difference from GLM-5.2 and DeepSeek V4, both of which are plain MIT.

Price and the cost-per-task trap

Moonshot prices K3 at $3 per million input tokens and $15 per million output, with cached input at $0.30 per million. That is roughly seven times the input price and seventeen times the output price of DeepSeek V4 Pro, and a large multiple of GLM-5.2.

Per-token price is the wrong unit for agent work. What you actually pay is tokens-per-task multiplied by price-per-token, and a model that finishes in four turns at $15/M output can beat one that flails for eighteen turns at $0.87/M. The Artificial Analysis cost-per-task figure of $0.94 is the more useful number, and it is competitive with frontier closed models.

But that arithmetic is workload-specific. On short, well-specified tasks where a cheaper model succeeds first time, K3 is straightforwardly more expensive with nothing to show for it. The 90% cache discount matters here: if you send a large stable system prompt or repository context on every call, most of your input cost collapses.

Self-hosting is mostly theoretical

Open weights are only useful if you can serve them. 1.56TB of parameters means a multi-node deployment with high-bandwidth interconnect before you serve a single request, and the quantised formats in the release help with memory but not with the fundamental node count.

For nearly every team the practical meaning of "open weights" here is provider choice and price competition, not running it in your own rack. Several providers serve K3 at rates matching Moonshot's. That is still valuable — it is the difference between one vendor and a market — but do not confuse it with local deployment.

How to decide whether it is worth it

Run the comparison on cost per completed task, not per token. Take ten real tasks from your git history, run K3 and a cheap open model such as DeepSeek V4 Pro through the same harness, and record three things: pass or fail, turns to completion, and total spend.

The likely result is that K3 wins on the hardest quarter of your workload and loses on the rest. That is a routing decision, not a replacement decision — and it is a far more useful finding than knowing which model tops a chart.

Common questions

Is Kimi K3 open source?

No. It is open weight. The weights are downloadable, but they ship under a custom kimi-k3 licence rather than MIT or Apache, and Model-as-a-Service businesses above $20 million in revenue over any twelve-month period need a separate agreement with Moonshot.

Can I run Kimi K3 on my own hardware?

Realistically, no. The weights are about 1.56TB, which means a multi-node cluster with fast interconnect. In practice the value of the open release is that multiple providers can serve it competitively, not that you will host it yourself.

How does Kimi K3 compare to closed frontier models?

It ranked first on Arena Frontend Code at 1,679 points, ahead of Claude Fable 5, and second only to Fable 5 on the Artificial Analysis long-horizon knowledge-work Elo. Moonshot acknowledges it trails Fable 5 and GPT-5.6 Sol overall and has a noticeable UX gap.

Similar articles

Kimi K3 vs GLM-5.2: Capability Against Deployability
Models
Models·9 min read

Kimi K3 vs GLM-5.2: Capability Against Deployability

One is the most capable open-weight model and hard to serve. The other is MIT-licensed, leaner and built for long agent loops. How to choose between them.

Read
DeepSeek V4 Pro: Frontier Reasoning at Commodity Prices
Models
Models·10 min read

DeepSeek V4 Pro: Frontier Reasoning at Commodity Prices

A 1.6T MoE that activates 49B per token, ships under MIT, and costs $0.435 per million input tokens. Where V4 Pro is strong, where it is not, and Pro versus Flash.

Read
Kimi K2.6 vs DeepSeek V4 Flash: Four Axes That Decide It
Models
Models·9 min read

Kimi K2.6 vs DeepSeek V4 Flash: Four Axes That Decide It

Two models released three days apart with opposite design goals. Active parameters, context length, licensing and vision decide which one fits your workload.

Read