Qwen 3.5 for Coding: Picking a Size You Can Actually Run
The Qwen line is the only open family that spans 27B to 397B under Apache 2.0. A guide to choosing a size, reading its benchmarks, and self-hosting economics.
Kimi K3, GLM-5.2 and DeepSeek V4 Pro all give you open weights you almost certainly cannot serve. At 1.56TB, 744B and 1.6T parameters respectively, "open" in practice means provider competition rather than local deployment.
Alibaba's Qwen line is the exception, and that is the whole reason to care about it. The Qwen3.5 family, released in February 2026 under Apache 2.0, spans sizes from a 27B dense model up to a 397B mixture-of-experts flagship, with 35B-A3B and 122B-A10B in between. Several of those genuinely fit on hardware you can rent by the hour or buy outright.
Read the version numbers before you read the benchmarks
This is the practical trap with Qwen, and it invalidates a lot of comparison posts. Alibaba ships fast and the naming overlaps: Qwen3-Coder (the 480B-A35B and 30B-A3B families, plus Qwen3-Coder-Next), the Qwen3.5 general line from February 2026, and later Qwen3.6 releases including a 35B-A3B under Apache 2.0.
Scores get attributed across these lines carelessly. Before you copy a number into a decision document, confirm three things: the exact model identifier including the active-parameter suffix, the benchmark variant (SWE-bench Verified and SWE-bench Pro are not interchangeable, and the Pro numbers are much lower), and whether the score came from the lab's own scaffold or a standardised harness.
That last one is not pedantry. Labs tune their agent scaffold alongside the model, and the same weights routinely score several points lower under a neutral harness. When Alibaba reports a SWE-bench Verified figure on its own scaffold, that is an upper bound on what you will see, not an estimate.
The architecture is built for long context on small hardware
The Qwen3.5 line uses a Gated DeltaNet plus MoE design that alternates linear and full attention in a 3:1 ratio. The purpose is near-linear compute scaling with context length, which is what makes very long windows affordable on modest hardware rather than merely representable.
The flagship Qwen3.5-397B-A17B carries 397B total parameters with 17B active per token, across 512 experts and 60 layers, with a native 262K context. The smaller members — 122B-A10B, 35B-A3B and the 27B dense model — apply the same recipe at sizes that a single node, or in the smallest cases a single high-memory GPU, can hold.
Reported results in this range are respectable rather than frontier: figures around 72.4 on SWE-bench Verified for Qwen3.5-27B and around 69 to 70 for 35B-A3B have circulated, with 122B-A10B posting 72.2 on BFCL-V4 for tool use against 55.5 for GPT-5 mini. Treat the exact digits as indicative — the sourcing varies — but the shape is consistent: these models land in the same band as small closed models, at a fraction of the marginal cost once you own the hardware.
Apache 2.0 is the strongest licence in the field
The open-weight Qwen models ship under Apache 2.0. That is a stronger position than every model it competes with. GLM-5.2 and DeepSeek V4 Pro are MIT, which is comparable. Kimi K3 uses a custom licence that gates Model-as-a-Service businesses above $20 million in annual revenue. MiniMax M3 uses a custom community licence.
Apache 2.0 adds an explicit patent grant on top of permissive redistribution, which is the clause corporate legal teams actually ask about. If you are fine-tuning weights and shipping the result inside a product, this is the licence you want.
The self-hosting arithmetic
Self-hosting only wins on sustained, predictable volume. The rough shape of the calculation:
- Measure your real token throughput over a normal week — input and output separately, since output is where API pricing hurts.
- Price that volume at API rates. DeepSeek V4-Flash at roughly $0.14 and $0.28 per million is a reasonable floor to beat; if you cannot beat it, stop here.
- Cost the hardware honestly. Amortised GPU cost, power, and — the line everyone omits — the engineer-days spent on serving infrastructure, batching, and upgrades.
- Divide by realistic utilisation. A cluster idle sixteen hours a day costs the same as a busy one. Bursty interactive workloads are the worst case for owned hardware.
The honest conclusion for most teams is that hosted APIs win on cost, and self-hosting wins on the things cost does not capture: data never leaving your network, no rate limits, no model deprecation on someone else's schedule, and the ability to fine-tune on proprietary code. Those are real reasons. Cost usually is not.
Choosing a size
Work down, not up. Start with the smallest model in the family and find where it breaks:
- 27B to 35B class — autocomplete, docstrings, test scaffolding, renames, commit messages, structured extraction. This handles more of a normal day than most people expect.
- ~122B class — multi-file feature work and tool calling, where instruction adherence and reliable function arguments matter more than raw reasoning.
- ~397B flagship — the tasks that survive the two tiers above.
Route rather than replace. The measurement that decides it is not an aggregate benchmark score but the failure rate of the small model on your own tasks, because that is the number that tells you how often you need to escalate — and therefore what the whole arrangement costs.
Common questions
Which Qwen model should I self-host for coding?
Start with the smallest size that fits your hardware, typically the 27B or 35B-A3B class, and measure its failure rate on your own tasks. Escalate to a larger member of the family only for the tasks it actually fails.
Why do Qwen benchmark numbers vary so much between articles?
Because the naming overlaps across Qwen3-Coder, Qwen3.5 and Qwen3.6, and because labs report on their own agent scaffolds. Always confirm the exact model identifier, the benchmark variant, and whether the harness was standardised.
Is self-hosting Qwen cheaper than using a hosted API?
Usually not, once you count idle time and engineering effort. Hosted open models are priced aggressively. Self-host for data residency, rate-limit freedom, version control and fine-tuning on proprietary code — not to save money.