MiniMax M3: Multimodal Agent Work From a 428B Open Model
Models

MiniMax M3: Multimodal Agent Work From a 428B Open Model

M3 is the smallest of the frontier-class open models and the only one that takes video natively. A look at sparse attention, the licence, and where it fits.

MiniMax M3 is the odd one out in the 2026 open-weight field. Every other model in the conversation is text-first with vision added later. M3 was trained multimodal from the start — text, image and video in, text out — and it is by far the smallest of the group at 428B total parameters with roughly 23B active per token.

The Shanghai lab released it on 1 June 2026. The pitch is that you do not need a 1.6T or 2.8T model to do frontier agentic coding, and that bundling multimodality into the same weights removes an orchestration problem rather than adding a feature.

Sparse attention is the design decision that matters

M3 uses a grouped-query attention backbone with MiniMax Sparse Attention, and supports a context window of up to 1M tokens. The point of the sparse attention scheme is to make long contexts cheap enough to actually use rather than merely advertise.

Around 23B active parameters is the leanest forward pass among the frontier open models — GLM-5.2 activates about 40B, DeepSeek V4 Pro 49B, Kimi K3 roughly 104B. That shows up directly in latency. Artificial Analysis measures M3 generating at about 112 tokens per second with a time to first token of roughly 1.5 seconds, faster on both counts than DeepSeek V4 Pro at around 74 tokens per second and 1.8 seconds.

In an agent loop, throughput compounds. A task that takes forty model calls at fifty extra tokens per second of throughput is a materially shorter wall-clock session, and interactive tools feel different at 1.5 seconds to first token than at 2.5.

The published benchmarks, read carefully

MiniMax published a focused set on release:

  • SWE-bench Pro — 59.0%
  • Terminal-Bench 2.1 — 66.0%
  • MCP-Atlas — 74.2%
  • SWE-fficiency — 34.8%
  • KernelBench Hard — 28.8%

They also cite a CUDA optimisation result on Hopper FP8 kernels, taking hardware utilisation from 7.6% to 71.3% — a 9.4x speedup on the generated kernel.

Compare that honestly against GLM-5.2, which reports 62.1 on SWE-bench Pro and 81.0 on Terminal-Bench 2.1. M3 is in the same conversation on repository-scale bug fixing and clearly behind on long-horizon terminal work. On the Artificial Analysis Intelligence Index it sits in the mid-forties, effectively tied with DeepSeek V4 Pro and below GLM-5.2 and Kimi K3.

Third-party write-ups circulate a figure of 80.5% on SWE-bench Verified. That number does not appear in MiniMax's own release material, so treat it as unconfirmed unless you find the harness it came from. SWE-bench Verified and SWE-bench Pro are different evaluations with very different difficulty, and conflating them is the most common error in model comparison posts.

Multimodality is a workflow argument, not a benchmark argument

The reason to care about native image and video input is not that it scores better. It is that some agent loops are much shorter when the model can see.

Concrete cases: a front-end agent that reads a screenshot of the broken layout rather than a DOM dump; a test-failure loop where the artefact is a visual diff; an agent triaging a bug report that arrives as a screen recording. Without native video input, each of these needs a separate captioning or vision model, an extra hop, and a lossy handoff.

Whether that is worth anything depends entirely on whether your inputs are visual. For a backend team working on a Go service it is dead weight. For anyone doing UI work or processing user-submitted media it removes a whole component from the pipeline.

Licence and price

The weights are on Hugging Face under a custom minimax-community licence. That is a weaker position than GLM-5.2 or DeepSeek V4 Pro, both plain MIT, and a stronger one than nothing. If you intend to serve M3 commercially, read the terms rather than assuming they behave like Apache.

API pricing is tiered on input length. Rates around $0.24 per million input tokens and $0.96 per million output have been listed under promotional discounting, against a standard rate closer to $0.60 and $2.40, with a higher band above 512K input tokens. Because the tiering keys on prompt length, a long-context workload can cost meaningfully more per token than the headline suggests — check which band your median request falls into before extrapolating a monthly bill.

Who should actually pick it

M3 makes sense when at least two of these are true: your inputs include images or video, latency matters because a human is waiting, and your agent tasks are medium-length rather than marathon sessions.

It makes less sense if your work is text-only long-horizon engineering, where GLM-5.2's Terminal-Bench profile is a better match, or if you are optimising purely for cost per token, where DeepSeek V4-Flash undercuts it.

The test worth running is the one that isolates the multimodal claim. Take five real bugs that arrived with a screenshot. Run them through M3 with the image attached, and through a text-only model with the same bug described in words. If the visual path does not reduce turns to completion, the multimodality is not buying you anything and you should choose on price and speed instead.

Common questions

How does MiniMax M3 compare to GLM-5.2 on coding?

M3 reports 59.0% on SWE-bench Pro and 66.0% on Terminal-Bench 2.1; GLM-5.2 reports 62.1% and 81.0% on the same two. They are close on repository-scale bug fixing, and GLM-5.2 is clearly ahead on long-horizon terminal work.

Is MiniMax M3 free to use commercially?

The weights are published under a custom minimax-community licence rather than MIT or Apache 2.0. Read the terms before building a commercial service on it — it is open weight, not an OSI-approved open-source licence.

Does native video input actually matter?

Only if your inputs are visual. It removes a captioning hop for screenshot-driven debugging, UI work and user-submitted media. For a text-only backend workload it buys nothing, and you should choose on speed and price instead.

Similar articles

DeepSeek V4 Flash vs MiniMax M3: Cheapest Against Multimodal
Models
Models·8 min read

DeepSeek V4 Flash vs MiniMax M3: Cheapest Against Multimodal

Two budget models with 1M context. One is cheaper and text-only under MIT, the other sees images. The choice is almost entirely about input type.

Read
GLM-5.2: The Open Model Built for Long-Horizon Coding
Models
Models·10 min read

GLM-5.2: The Open Model Built for Long-Horizon Coding

Z.ai shipped a 744B MoE with 40B active, MIT-licensed weights and the first open-weight Terminal-Bench 2.1 score above 80. A technical read on what that means.

Read
Kimi K2.6 vs MiniMax M3: Two Ways to Do Multimodal
Models
Models·9 min read

Kimi K2.6 vs MiniMax M3: Two Ways to Do Multimodal

Both take images natively, but one gives you 256K of context at a known price and the other 1M at a price sources disagree on. How to pick between them.

Read