FW-MiniMax-M3
Models available for use with Fireworks on Foundry deliver optimized, best-in-class performance on the Fireworks Inference Cloud. Fireworks on Foundry is a Non-Microsoft Product. The following terms apply to a Customer's use of Fireworks on Foundry: When you use Fireworks on Foundry, data is shared between Microsoft and Fireworks AI, Customer Data will be sent outside of Microsoft systems, Customer Data will not be processed pursuant to any Foundry data residency documentation, and different compliance and data handling rules will apply. See Trust Center - Fireworks AI for details. Customers are responsible for evaluating whether data sharing between Microsoft and Fireworks is appropriate for their organization's compliance requirements.
About this model
MiniMax-M3 is a native multimodal model from MiniMax with 1M context, about 428B parameters, and about 23B activated parameters. It undergoes mixed-modality training from the first step, enabling semantic fusion across text and image inputs. The model introduces MiniMax Sparse Attention (MSA), a sparse attention operator designed for million-token contexts, and targets frontier-level performance across long-horizon agentic benchmarks, coding, and cowork tasks.
Key model capabilities
- Native multimodality across text and image inputs
- 1M token context length
- MiniMax Sparse Attention for efficient long-context processing
- Mixture-of-Experts architecture with about 428B total parameters and about 23B activated parameters
- Coding and cowork capabilities for long-horizon agentic benchmarks
- Reasoning modes controlled by the
thinkingparameter: enabled, adaptive, and disabled