Open Weights vs Open Source: The Distinction That Matters
Downloadable weights are not the same thing as open source. Here is what the licences actually permit, and which differences change what you can ship.
"Open source model" is used to mean at least three different things: weights you can download, weights under a licence approved by the Open Source Initiative, and a full training pipeline you could reproduce. Only the first is common.
The distinction is not pedantry. It decides whether you can self-host, fine-tune, redistribute, or build a competing product without a lawyer in the room.
What open weights actually gives you
Open weights means the trained parameters are downloadable. You can run inference on your own hardware, quantise them, fine-tune them, and inspect activations. That covers most of what engineers care about day to day.
It does not usually include the training data, the data pipeline, the training code, or the recipe. You can modify the model by continuing to train it, but you cannot rebuild it. In software terms it is closer to a redistributable binary than to source.
What the OSI definition requires
The Open Source Initiative published version 1.0 of the Open Source AI Definition in late 2024 to settle exactly this argument. It asks for the four familiar freedoms — use, study, modify, share — applied to an AI system, which in practice means the weights, the code needed to run and train, and sufficient information about the data for a skilled person to build a substantially equivalent system.
Critically, the freedom to use must be unconditional. Any use-case restriction, user-class restriction or field-of-endeavour limitation disqualifies a licence regardless of how permissive the rest of it reads.
That single clause is what most model licences fail.
The three licence families
Genuinely permissive. MIT and Apache 2.0 releases impose no use restrictions and no revenue thresholds. GLM 5.2 from Z.ai and DeepSeek V4 Pro both ship under MIT, which means commercial use, fine-tuning and redistribution are all uncontroversial. If you want to avoid legal review entirely, this is the family to shortlist.
Bespoke community licences. These read like open source until you reach the conditions. Meta's Llama Community License carries an acceptable use policy and a threshold requiring a separate agreement above 700 million monthly active users. The OSI has stated plainly that this is not open source, on the grounds that it restricts the freedom to use and discriminates between classes of user.
Kimi K3 landed in this family too. Moonshot published the 2.8T-parameter weights in late July 2026, but under a bespoke Kimi K3 document rather than the modified MIT terms that shipped with K2 — with a revenue-triggered separate-agreement clause for model-as-a-service operators and an attribution requirement above a large monthly-active-user threshold. Open weights, conditional commercial terms.
Weights-available research licences. Non-commercial or research-only terms. Fine for evaluation, unusable for a product.
Why the difference is practical, not ideological
- Redistribution. Shipping a fine-tuned model to customers is straightforward under MIT and a legal question under a community licence.
- Serving as a service. Revenue and user thresholds usually bite exactly when you succeed, and are easy to forget at the point you designed around them.
- Naming and attribution. Several bespoke licences require derivative model names or UI attribution to reference the base model.
- Acceptable use policies. These bind your users, not just you, and someone has to enforce them.
- Change of terms. A permissive licence on a downloaded artefact cannot be revoked for that version. A bespoke one may carry terms that change for future releases.
What you gain by using open weights at all
Set the licence question aside for a moment — even a restrictive open-weights model buys you things a closed API cannot.
Portability is the big one. A model you can download cannot be deprecated out from under you. When a hosted endpoint is retired, your prompts, your evaluation results and your tuned behaviour go with it. When you hold the weights, the worst case is that you keep serving the version you already have.
Price competition follows from the same property. Because multiple providers can serve identical weights, open-weight models are subject to real price pressure in a way that a single-vendor endpoint is not. You can move between hosts without re-validating model behaviour, because it is the same model.
And self-hosting becomes possible for the cases that require it: data residency, air-gapped environments, or regulated workloads where sending text to a third party is not an option. Possible is not the same as cheap — serving a large mixture-of-experts model well needs serious hardware and real expertise — but the option exists.
What only true open source gives you
Reproducibility and auditability. Without data information you cannot check what a model was trained on, cannot verify claims about contamination or provenance, and cannot rebuild it if the weights disappear.
For most product teams that is an acceptable gap. For research, safety work, or procurement in regulated sectors where "what is in this model" is an answerable requirement, it is the whole point.
How to read a model licence in five minutes
- Search for a revenue or user threshold. If there is one, note the number and where you sit against it.
- Look for an acceptable use policy and check whether you are required to pass it through to your users.
- Check whether redistribution of a fine-tuned derivative is permitted, and under what name.
- Check whether outputs may be used to train other models. Some licences forbid it.
- Confirm which artefacts are covered — weights only, or code and tokeniser too.
Then apply the decision rule. If you are evaluating, almost anything works. If you are shipping a product, prefer MIT or Apache and skip the review. If a bespoke-licence model is meaningfully better for your task, read the thresholds carefully and write down where you would cross them — that is a business decision, and it should be made deliberately rather than discovered later.
Common questions
Is Llama open source?
Not by the OSI definition. The weights are downloadable, but the Llama Community License adds an acceptable use policy and a monthly-active-user threshold, and the OSI has said publicly that it is not open source.
Which open models have genuinely permissive licences?
GLM 5.2 and DeepSeek V4 Pro both ship under MIT, which imposes no use restrictions or revenue thresholds. Always confirm on the model card, since terms sometimes change between releases.
Does open weights mean I can self-host cheaply?
It means you are allowed to. Serving a large mixture-of-experts model at good latency needs substantial GPU memory and tuning, so hosted inference is often cheaper below high, steady volume.