MiniMax M3: Multimodal Agent Work From a 428B Open Model
M3 is the smallest of the frontier-class open models and the only one that takes video natively. A look at sparse attention, the licence, and where it fits.
MiniMax M3 is the odd one out in the 2026 open-weight field. Every other model in the conversation is text-first with vision added later. M3 was trained multimodal from the start — text, image and video in, text out — and it is by far the smallest of the group at 428B total parameters with roughly 23B active per token.
The Shanghai lab released it on 1 June 2026. The pitch is that you do not need a 1.6T or 2.8T model to do frontier agentic coding, and that bundling multimodality into the same weights removes an orchestration problem rather than adding a feature.
Sparse attention is the design decision that matters
M3 uses a grouped-query attention backbone with MiniMax Sparse Attention, and supports a context window of up to 1M tokens. The point of the sparse attention scheme is to make long contexts cheap enough to actually use rather than merely advertise.
Around 23B active parameters is the leanest forward pass among the frontier open models — GLM-5.2 activates about 40B, DeepSeek V4 Pro 49B, Kimi K3 roughly 104B. That shows up directly in latency. Artificial Analysis measures M3 generating at about 112 tokens per second with a time to first token of roughly 1.5 seconds, faster on both counts than DeepSeek V4 Pro at around 74 tokens per second and 1.8 seconds.
In an agent loop, throughput compounds. A task that takes forty model calls at fifty extra tokens per second of throughput is a materially shorter wall-clock session, and interactive tools feel different at 1.5 seconds to first token than at 2.5.
The published benchmarks, read carefully
MiniMax published a focused set on release:
- SWE-bench Pro — 59.0%
- Terminal-Bench 2.1 — 66.0%
- MCP-Atlas — 74.2%
- SWE-fficiency — 34.8%
- KernelBench Hard — 28.8%
They also cite a CUDA optimisation result on Hopper FP8 kernels, taking hardware utilisation from 7.6% to 71.3% — a 9.4x speedup on the generated kernel.
Compare that honestly against GLM-5.2, which reports 62.1 on SWE-bench Pro and 81.0 on Terminal-Bench 2.1. M3 is in the same conversation on repository-scale bug fixing and clearly behind on long-horizon terminal work. On the Artificial Analysis Intelligence Index it sits in the mid-forties, effectively tied with DeepSeek V4 Pro and below GLM-5.2 and Kimi K3.
Third-party write-ups circulate a figure of 80.5% on SWE-bench Verified. That number does not appear in MiniMax's own release material, so treat it as unconfirmed unless you find the harness it came from. SWE-bench Verified and SWE-bench Pro are different evaluations with very different difficulty, and conflating them is the most common error in model comparison posts.
Multimodality is a workflow argument, not a benchmark argument
The reason to care about native image and video input is not that it scores better. It is that some agent loops are much shorter when the model can see.
Concrete cases: a front-end agent that reads a screenshot of the broken layout rather than a DOM dump; a test-failure loop where the artefact is a visual diff; an agent triaging a bug report that arrives as a screen recording. Without native video input, each of these needs a separate captioning or vision model, an extra hop, and a lossy handoff.
Whether that is worth anything depends entirely on whether your inputs are visual. For a backend team working on a Go service it is dead weight. For anyone doing UI work or processing user-submitted media it removes a whole component from the pipeline.
Licence and price
The weights are on Hugging Face under a custom minimax-community licence. That is a weaker position than GLM-5.2 or DeepSeek V4 Pro, both plain MIT, and a stronger one than nothing. If you intend to serve M3 commercially, read the terms rather than assuming they behave like Apache.
API pricing is tiered on input length. Rates around $0.24 per million input tokens and $0.96 per million output have been listed under promotional discounting, against a standard rate closer to $0.60 and $2.40, with a higher band above 512K input tokens. Because the tiering keys on prompt length, a long-context workload can cost meaningfully more per token than the headline suggests — check which band your median request falls into before extrapolating a monthly bill.
Who should actually pick it
M3 makes sense when at least two of these are true: your inputs include images or video, latency matters because a human is waiting, and your agent tasks are medium-length rather than marathon sessions.
It makes less sense if your work is text-only long-horizon engineering, where GLM-5.2's Terminal-Bench profile is a better match, or if you are optimising purely for cost per token, where DeepSeek V4-Flash undercuts it.
The test worth running is the one that isolates the multimodal claim. Take five real bugs that arrived with a screenshot. Run them through M3 with the image attached, and through a text-only model with the same bug described in words. If the visual path does not reduce turns to completion, the multimodality is not buying you anything and you should choose on price and speed instead.
Common questions
How does MiniMax M3 compare to GLM-5.2 on coding?
M3 reports 59.0% on SWE-bench Pro and 66.0% on Terminal-Bench 2.1; GLM-5.2 reports 62.1% and 81.0% on the same two. They are close on repository-scale bug fixing, and GLM-5.2 is clearly ahead on long-horizon terminal work.
Is MiniMax M3 free to use commercially?
The weights are published under a custom minimax-community licence rather than MIT or Apache 2.0. Read the terms before building a commercial service on it — it is open weight, not an OSI-approved open-source licence.
Does native video input actually matter?
Only if your inputs are visual. It removes a captioning hop for screenshot-driven debugging, UI work and user-submitted media. For a text-only backend workload it buys nothing, and you should choose on speed and price instead.