DeepSeek V4 Pro: Frontier Reasoning at Commodity Prices
Models

DeepSeek V4 Pro: Frontier Reasoning at Commodity Prices

A 1.6T MoE that activates 49B per token, ships under MIT, and costs $0.435 per million input tokens. Where V4 Pro is strong, where it is not, and Pro versus Flash.

DeepSeek V4 Pro is the model that makes the rest of the market explain itself. It is a 1.6 trillion parameter mixture-of-experts model with 49 billion active parameters, a 1M-token context window, MIT-licensed weights, and API pricing of roughly $0.435 per million input tokens and $0.87 per million output — a permanent 75% cut applied at the end of May 2026.

That is roughly one seventh the input price of Kimi K3 and one seventeenth the output price. The interesting question is not whether it is cheap. It is what you give up, and the answer is more specific than "quality".

The architecture, and why the price is possible

V4 Pro activates 49B of its 1.6T parameters per forward pass — an activation ratio of about 3%. That sparsity is the whole story of the price. You need enormous memory to hold the model, but the compute spent generating each token is closer to a mid-sized dense model than to a frontier one.

DeepSeek pairs this with a hybrid attention system for long-context handling, which is what makes the 1M window practical rather than nominal. Maximum output is 384K tokens, and the model exposes high and extended reasoning effort levels.

Artificial Analysis measures output at about 73.7 tokens per second with a time to first token of 1.81 seconds, and reports a blended price of roughly $0.18 per million tokens on a 7:2:1 cache-hit / input / output mix. Whether that blend resembles your workload depends entirely on how much of your context repeats.

Where it is genuinely strong

DeepSeek's official V4 report puts V4-Pro-Max at 80.6% on SWE-bench Verified, 93.5 on LiveCodeBench, 87.5 on MMLU-Pro, 90.1 on GPQA Diamond, 95.2 pass@1 on AIME 2025, and a Codeforces rating of 3206.

Those are vendor-run figures, not independent leaderboard entries, and vendor scaffolds reliably score above standardised harnesses. Read them as a shape rather than as absolutes — and the shape is consistent across sources: this model is disproportionately good at self-contained algorithmic and mathematical problems. A Codeforces rating in the 3200s and a 93.5 LiveCodeBench are the signature of a model trained hard on problems with a verifiable answer.

On the Artificial Analysis Intelligence Index it lands in the mid-forties, which places it alongside MiniMax M3 among open models and below GLM-5.2 and Kimi K3. That gap is the honest summary: V4 Pro is not the most capable open model, it is the one with the best ratio of capability to price by a wide margin.

Where it is weaker

Reporting on V4 Pro consistently identifies two soft spots: pure factual recall, and hard multi-step reasoning where the path is not verifiable at each step. That is the mirror image of its strength. A model optimised on problems with checkable answers is not automatically good at judgement calls in an unfamiliar codebase.

The practical consequence for agent work is that the failure mode is confident and plausible rather than obviously broken. On an algorithm it either passes the tests or does not. On "figure out why this integration is flaky", a wrong hypothesis stated fluently costs you more review time than a visibly bad answer would.

It also does not lead on the benchmarks built specifically for long-horizon agentic coding — that is GLM-5.2's territory, and the difference shows up in tool-heavy sessions rather than in single-turn quality.

Pro versus Flash

DeepSeek ships V4 in two sizes and most teams should be running the smaller one for most work. V4-Flash is 284B total with 13B active, priced around $0.14 per million input and $0.28 per million output — roughly a third of Pro.

On practical coding the gap is narrow. Flash lands within a few points of Pro on SWE-bench Verified, and a July 2026 re-post-training pushed its Terminal-Bench 2.1 score to 82.7 at unchanged pricing. On the Intelligence Index there is a real but modest gap between them.

The decision rule that falls out: default to Flash for production code generation, retrieval pipelines and tool calling; escalate to Pro for deep agentic loops, hallucination-sensitive work, and problems that are genuinely hard rather than merely long. Sending everything to Pro is the same mistake as sending everything to a closed frontier model, just cheaper.

MIT weights, and what that is worth

V4 Pro is MIT-licensed. No revenue threshold, no separate agreement for serving it commercially, no regional clause. Among the frontier-adjacent open models only GLM-5.2 matches that; Kimi K3 uses a custom licence with a Model-as-a-Service revenue gate, and MiniMax M3 ships under a custom community licence.

Serving 1.6T parameters yourself is still a serious infrastructure commitment even at 49B active, so for most teams the licence buys provider competition rather than local deployment. That competition is visible in the price — multiple hosts, aggressive discounting, and no single vendor able to raise rates unilaterally.

A decision rule

Use V4 Pro as the baseline against which you justify anything more expensive. Concretely:

  1. Run your eval set on V4-Flash first. Record pass rate, turns to completion and total spend.
  2. Re-run the failures on V4 Pro. If Pro clears most of them, you have found your escalation tier.
  3. Re-run whatever still fails on GLM-5.2 or Kimi K3. That residue is what a premium model is genuinely for.

Teams that run this ladder usually discover the expensive tier is needed on a small minority of tasks. Knowing precisely which minority is worth more than any aggregate score.

Common questions

Why is DeepSeek V4 Pro so much cheaper than other frontier-class models?

Sparsity. It activates about 49B of 1.6T parameters per token, roughly a 3% activation ratio, so the compute spent generating each token is far below what the headline size suggests. DeepSeek also applied a permanent 75% price cut in May 2026.

Should I use V4 Pro or V4 Flash?

Flash for most production work — it is 284B total with 13B active at around a third of the price, and lands within a few points of Pro on practical coding. Escalate to Pro for deep agent loops and genuinely hard reasoning.

What is DeepSeek V4 Pro worst at?

Pure factual recall and hard multi-step reasoning where no step is independently verifiable. Its training profile favours problems with checkable answers, so ambiguous judgement calls in unfamiliar code are where it fails most confidently.

Similar articles

Kimi K3 vs DeepSeek V4 Pro: Doing the 17x Price Arithmetic
Models
Models·10 min read

Kimi K3 vs DeepSeek V4 Pro: Doing the 17x Price Arithmetic

K3 costs about seventeen times more per output token than DeepSeek V4 Pro. Here is the break-even calculation that decides whether that is worth paying.

Read
GLM-5.2 vs DeepSeek V4 Pro: Endurance or Raw Problem Solving
Models
Models·9 min read

GLM-5.2 vs DeepSeek V4 Pro: Endurance or Raw Problem Solving

Both are MIT-licensed open models, so the licence is a wash. The split is agentic endurance against algorithmic power, at a three-to-five times price gap.

Read
Kimi K2.6 vs DeepSeek V4 Flash: Four Axes That Decide It
Models
Models·9 min read

Kimi K2.6 vs DeepSeek V4 Flash: Four Axes That Decide It

Two models released three days apart with opposite design goals. Active parameters, context length, licensing and vision decide which one fits your workload.

Read