MiniMax M3 vs DeepSeek V4 Pro: Same Score, Different Models
Both land in the mid-forties on the Artificial Analysis index. They are not interchangeable, and the tie is a good lesson in why aggregate scores mislead.
MiniMax M3 and DeepSeek V4 Pro both land in the mid-forties on the Artificial Analysis Intelligence Index. If you were shopping from a leaderboard you would conclude they are interchangeable and pick the cheaper one.
They are not interchangeable. One is a 1.6T text reasoning model tuned on verifiable problems; the other is a 428B natively multimodal model that generates tokens 50% faster. The tie is a useful demonstration of what an aggregate index throws away — four benchmark categories averaged into one number will happily rate two very differently-shaped models the same.
The specifications diverge more than the score does
DeepSeek V4 Pro is 1.6T total parameters with 49B active, an activation ratio near 3%, released 24 April 2026 under MIT, with a 1M context window and 384K maximum output.
MiniMax M3 is 428B total with roughly 23B active, released 1 June 2026 under a custom minimax-community licence, also with up to 1M context. It uses a grouped-query attention backbone with MiniMax Sparse Attention, and takes text, image and video as input natively — trained that way from the beginning rather than adapted afterwards.
M3 is therefore about a quarter the total size and roughly half the active parameters. That is the leanest forward pass of any frontier-class open model, and it is visible in the measurements: Artificial Analysis clocks M3 at about 112 tokens per second output with a 1.5-second time to first token, against roughly 74 tokens per second and 1.8 seconds for V4 Pro.
What each lab chose to measure
MiniMax published a compact agentic set on release: 59.0% on SWE-bench Pro, 66.0% on Terminal-Bench 2.1, 74.2% on MCP-Atlas, 34.8% on SWE-fficiency and 28.8% on KernelBench Hard. They also cite a CUDA kernel result on Hopper FP8 taking hardware utilisation from 7.6% to 71.3%.
DeepSeek's V4 report emphasises something else: 80.6% on SWE-bench Verified for V4-Pro-Max, 93.5 on LiveCodeBench, 90.1 on GPQA Diamond, 95.2 pass@1 on AIME 2025 and a Codeforces rating of 3206.
A caution before you compare these directly — you cannot. SWE-bench Verified and SWE-bench Pro are different evaluations with substantially different difficulty, and the numbers are not on the same scale. Third-party posts that place M3 at 80.5% on SWE-bench Verified are not sourced to MiniMax's own release material, so treat that figure as unconfirmed until you find the harness behind it. This confusion is the single most common error in comparison articles about these two models.
What you can compare is emphasis. MiniMax measured tool use, terminal sessions and efficiency of the generated code. DeepSeek measured competitive programming, mathematics and graduate science. Labs benchmark what they trained for.
Latency is a capability, not a comfort
The speed gap deserves more weight than it usually gets. In an agent loop, throughput multiplies: a forty-turn task at 112 tokens per second finishes meaningfully sooner than the same task at 74, and the time-to-first-token difference is the one a human actually feels in an interactive tool.
This is where M3 earns its place despite the smaller parameter count. Sparse attention plus 23B active parameters buys responsiveness, and responsiveness changes how people use a tool — faster loops mean more iterations, and more iterations often beat a better single answer.
Against that, V4 Pro supports 384K output tokens and two reasoning effort levels, which suits tasks where you want the model to think longer rather than reply sooner.
Multimodality is the real fork
The honest decision rule between these two has little to do with the index score.
If your inputs are visual — screenshots of broken layouts, visual regression diffs, screen recordings attached to bug reports, user-submitted media — M3 handles them in the same weights. The alternative is a separate vision model, an extra hop, and a lossy handoff into text. That removed component is worth more than a couple of index points.
If your inputs are text and code, native video buys you nothing and you are paying for capability you never invoke. V4 Pro is the stronger text reasoner and, at roughly $0.435 per million input and $0.87 per million output, is priced against MiniMax M3 at about $0.24 and $0.96 under promotional rates, closer to $0.60 and $2.40 at standard rates with a higher band above 512K input tokens.
Note the shape of that: M3 can be cheaper on input and more expensive on output. Which model costs you less depends on your input-to-output ratio, so compute it from your actual logs rather than eyeballing the rate card.
Licence, and how to decide
V4 Pro is MIT — unrestricted commercial use, redistribution and fine-tuning. M3 uses a custom minimax-community licence, which is open weight but not an OSI-approved open-source licence. If you plan to serve either model to paying customers, read M3 terms properly first.
Then run the one experiment that separates them. Take five real bugs that arrived with a screenshot or recording. Give M3 the artefact directly, and give V4 Pro a careful text description of the same thing. Count turns to resolution.
If the visual path is shorter, that is your answer and no index score was needed. If it is not, you have a text workload, and you should choose on reasoning quality and cost per token instead.
Common questions
Why do MiniMax M3 and DeepSeek V4 Pro score the same if they are so different?
The Artificial Analysis index averages four weighted categories into one number. Two models can reach the same average through completely different strengths — M3 through agentic and multimodal work, V4 Pro through mathematics and algorithmic coding.
Which one is faster?
MiniMax M3, by a clear margin. Artificial Analysis measures roughly 112 tokens per second output and a 1.5-second time to first token, against about 74 tokens per second and 1.8 seconds for DeepSeek V4 Pro. M3 activates roughly 23B parameters per token against 49B.
Which is cheaper?
It depends on your input-to-output ratio. M3 can undercut V4 Pro on input under promotional pricing while costing more per output token, and M3 charges a higher band above 512K input tokens. Compute it from your own logs.