MAI-Thinking-1
Direct from Azure models are a select portfolio curated for their market-differentiated capabilities:
- Secure and managed by Microsoft: Purchase and manage models directly through Azure with a single license, consistent support, and no third-party dependencies, backed by Azure's enterprise-grade infrastructure.
- Streamlined operations: Benefit from unified billing, governance, and seamless PTU portability across models hosted on Azure - all as part of one Microsoft Foundry platform.
- Future-ready flexibility: Access the latest models as they become available, and easily test, deploy, or switch between them within Microsoft Foundry; reducing integration effort.
- Cost control and optimization: Scale on demand with pay-as-you-go flexibility or reserve PTUs for predictable performance and savings.
Learn more about Direct from Azure models .
About this model
MAI-Thinking-1 is a sparse Mixture-of-Experts Transformer with 35B active and 1T total parameters, developed by Microsoft AI. It is a reasoning model: given a prompt, it produces an internal chain of thought before emitting a final response, allocating reasoning effort adaptively according to prompt complexity.
Pre-training used a general-purpose text corpus spanning code, academic and PDF content, math and STEM web content, books, structured knowledge sources, and general web content. Post-training combined supervised fine-tuning on model-generated completions with reinforcement learning across verifiable reasoning, software engineering, tool-use, and instruction-following environments.
MAI-Thinking-1 focuses on reasoning efficiency and agentic capability rather than a change in architecture or scale. Chain-of-thought sequences used in post-training were length-compressed while preserving task performance, reducing the token cost of a typical response. Post-training additionally expanded coverage of tool use, agentic task completion, instruction following, factuality, and over-refusal reduction.
Key model capabilities
- Reasoning: Optimized for top-tier reasoning. Achieves state-of-the-art performance on math, knowledge, and coding for its weight class.
- Coding performance: Matches frontier-model performance on SWE-Bench Pro and trained using 8M+ RLE environment
- Price-to-performance: Adaptive reasoning that matches effort to prompt complexity, best price-to-performance ratio for reasoning and coding tasks.
- Enterprise ready: Built on top of clean data that is appropriately licensed, which allows for quality, provenance, and control.
- Seamless migration: Built on top of the widely used Chat Completions API, making migrations easy.
- Enterprise deployments: ready for enterprise deployments with a 256K context window, clean data provenance, function calling, and the ability to follow complex instructions.
- Coding workflows: reading code, editing files, running tests, bug fixing, observing failures, and recovering intermediate mistakes.
- Complex reasoning tasks: especially strong at tasks that require quantitative reasoning like financial modeling, statistical analysis, market sizing, forecasting, etc.
Out of scope use cases
MAI-Thinking-1 is not designed or evaluated for use as an autonomous decision-maker in consequential domains. It should not be used as the sole basis for decisions with legal, financial, medical, employment, educational, housing, credit, or safety-critical consequences, and it is not a substitute for professional advice in regulated fields. It is not evaluated for use in fully autonomous agentic deployments that act on untrusted external content without human oversight or harness-level controls. The model has no native tool interface, tool use is mediated entirely by the integrating application, and the integrator is responsible for the security boundary around any capability it exposes.
The model is text-only and does not accept or produce image, audio, or video content. Uses prohibited by the applicable acceptable use policy and product terms are out of scope, including generation of disallowed content and use in applications designed to deceive, surveil, or manipulate individuals. Performance in languages other than those listed in §2.5 has not been systematically evaluated and such use is out of scope.