Skip to main content
Microsoft Foundry
Microsoft AI

MAI-Voice-2.1

MAI-Voice-2.1 is a prompted text-to-speech (TTS) model that generates high-fidelity, natural, and expressive speech across 10+ languages. It captures human-like intonation, rhythm, and emotional nuance for engaging conversational experiences.
Microsoft
Version: 2026-09-19

MAI-Voice-2.1 is a prompted text-to-speech (TTS) model built in-house by the Microsoft AI team, generating high-fidelity, natural, and expressive speech across 23 languages. It captures human-like intonation, rhythm, and emotional nuance, delivers emotional flexibility with turn-level control over tone and delivery, and can render a single voice identity consistently across every supported language, making it ideal for audiobooks, content creation, voice-over, media, and other scenarios where fidelity and expressiveness matter most.

Voice can be configured using:

  • Curated voice library (licensed voices designed to work straight out of the box)
  • Voice Cloning via short audio clips (5-60 seconds), and the model matches the voice instantly

  • Natural and expressive voice synthesis with realistic intonation, rhythm, and emotional range.
  • High-fidelity, high-clarity voice output.
  • Multilingual support across 23 languages.
  • One voice across all supported languages: a single voice identity is rendered consistently in every supported language.
  • Fine-grained control of tone and delivery: shape emotion, tone, and delivery at the turn/sentence level, including different speaking styles via SSML.
  • State-of-the-art instant Voice Cloning: provide a short audio clip (5-60 seconds) and the model clones the voice instantly, with no fine-tuning required. Prompt quality significantly impacts output, with best results from natural, conversational delivery and moderate energy levels. Access requires Microsoft approval and guardrails are in place to avoid misuse.
  • Long-form content generation with improved consistency via chunking and context carryover.
  • Multi-speaker support for dialogue and conversational scenarios.

Quick facts

PublisherMicrosoft
AuthorMicrosoft AI
TypeText to speech, Audio generation
LifecyclePreview
Input typetext
Output typeaudio