MiniMax M3: judge it by the 24 hour run
MiniMax's new flagship pairs a 1M token context with the metric that matters for agentic delivery: how long it stays coherent without a human.
MiniMax M3 is the company’s new flagship: frontier level coding and agentic performance, a 1 million token context window on its sparse attention architecture, and native multimodal understanding in one model. The showcase worth an executive’s attention is not a chat transcript. MiniMax had M3 optimize an FP8 GEMM kernel on Nvidia Hopper GPUs: roughly 24 hours, 147 benchmark submissions, 1,959 tool calls, hardware utilization pushed from 7.6 percent to 71.3 percent, no human in the loop.
That is the property agentic delivery actually depends on. Most models can produce a plausible first hour of work. Very few stay coherent through hour twenty. If you are evaluating models for workflow automation or agentic delivery stages, long horizon stability should be a first class procurement criterion, tested on your own tasks rather than taken from a launch page.
Two caveats. The open weights are promised on Hugging Face but not published yet, so treat self-hosting plans as futures. And MiniMax retired M2.1 in mid July, about seven months after it shipped. That deprecation pace is normal for this tier now: pin versions, and budget a migration lane as a permanent line item.
Written by Adib Kadir. Product and engineering executive focused on rolling out AI at enterprise scale.
Start a conversation