GLM-5.2 vs Qwen 3.6: Sparse and Huge, or Dense and Small
Both are permissively licensed and strong at code. One is a 744B mixture-of-experts, the other a dense 27B on a single GPU. The architecture is the decision.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Picking between models — and testing them on your own work.
Both are permissively licensed and strong at code. One is a 744B mixture-of-experts, the other a dense 27B on a single GPU. The architecture is the decision.
ReadModel cards mix hard facts, marketing and careful omissions. Which fields are reliable, which need checking, and what an absent section tells you.
ReadMoonshot shipped K3 four months after K2.6, but the older model did not become useless. What K2.6 is, what it costs, and the jobs it is still the right pick for.
ReadTwo models released three days apart with opposite design goals. Active parameters, context length, licensing and vision decide which one fits your workload.
ReadK2.6 sees images at 256K context under a custom licence. V4 Pro is text-only, MIT, 1M context, and reasons harder. Two models that barely overlap.
ReadK2.6 sees images and stops at 256K context under a custom licence. GLM-5.2 is text-only, MIT, and takes 1M tokens. Two clean trade-offs, no overlap.
ReadBoth take images natively, but one gives you 256K of context at a known price and the other 1M at a price sources disagree on. How to pick between them.
ReadA trillion-parameter vision model you rent against a dense 27B you can own. The comparison is about deployment shape, not a few points of benchmark difference.
ReadThe strongest open-weight model against one of the cheapest, both with 1M context. When a twenty-fold price difference is worth paying and when it is waste.
ReadK3 is the stronger model. K2.6 costs roughly a quarter as much per output token and activates a third of the parameters. When the older model is still the right call.
ReadOne is the strongest open-weight model available. The other is among the cheapest capable ones and natively multimodal. The gap between them is a budget decision.
ReadK3 needs roughly 1.6TB of weights and multi-node serving. Qwen 3.6 27B is dense and fits on one accelerator. The comparison is about deployment, not benchmarks.
ReadShowing 49–60 of 119 articles