Kimi K3 vs DeepSeek V4 Pro: Doing the 17x Price Arithmetic
K3 costs about seventeen times more per output token than DeepSeek V4 Pro. Here is the break-even calculation that decides whether that is worth paying.
ReadPractical writing for developers building with large language models — how they work, how to pick one, and how to keep the bill predictable.
Picking between models — and testing them on your own work.
K3 costs about seventeen times more per output token than DeepSeek V4 Pro. Here is the break-even calculation that decides whether that is worth paying.
ReadOne is the most capable open-weight model and hard to serve. The other is MIT-licensed, leaner and built for long agent loops. How to choose between them.
ReadMoonshot shipped the largest open-weight model yet: 2.8T parameters, 1M context, and a licence that is not MIT. A technical read on where K3 earns its price.
ReadBoth land in the mid-forties on the Artificial Analysis index. They are not interchangeable, and the tie is a good lesson in why aggregate scores mislead.
ReadM3 is the smallest of the frontier-class open models and the only one that takes video natively. A look at sparse attention, the licence, and where it fits.
ReadMost requests do not need your most capable model. Routing by task cuts cost sharply and adds resilience — here is how to build it without a mess.
ReadKimi K3, GLM-5.2, DeepSeek V4, MiniMax M3 and Qwen compared on the axes that decide deployment: licence terms, active parameters, serving cost and benchmark shape.
ReadThe Qwen line is the only open family that spans 27B to 397B under Apache 2.0. A guide to choosing a size, reading its benchmarks, and self-hosting economics.
ReadMixture-of-experts split model size into total and active parameters, and only one of them predicts your bill. How to think about size when picking a model.
ReadPublic leaderboards tell you almost nothing about your codebase. Here is how to build a private eval set in an afternoon and actually know which model to use.
ReadRefactoring is a constraint problem, not a generation problem. What the task actually demands from a model, and how to measure it on your own repo.
ReadShowing 109–119 of 119 articles