Back to Catalog

MiniMax-M3

Model

MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters. It undergoes mixed-modality training from the very first step, enabling deeper semantic fusion across text, image, and video. M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M context, reducing per-token compute to 1/20.

Contributors
MiniMax
Language
sk, eng, other
Modality
Other