Back to Catalog
MiniMax-M3
Model
MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters. It undergoes mixed-modality training from the very first step, enabling deeper semantic fusion across text, image, and video. M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M context, reducing per-token compute to 1/20.
- Contributors
- MiniMax
- Language
- sk, eng, other
- Modality
- Other