MAI-Voice-2.1
MAI-Voice-2.1 is a prompted text-to-speech (TTS) model that generates high-fidelity, natural, and expressive speech across 10+ languages. It captures human-like intonation, rhythm, and emotional nuance for engaging conversational experiences.
MAI-Voice-2.1 is a prompted text-to-speech (TTS) model built in-house by the Microsoft AI team, generating high-fidelity, natural, and expressive speech across 23 languages. It captures human-like intonation, rhythm, and emotional nuance, delivers emotional flexibility with turn-level control over tone and delivery, and can render a single voice identity consistently across every supported language, making it ideal for audiobooks, content creation, voice-over, media, and other scenarios where fidelity and expressiveness matter most.
Voice can be configured using:
- Curated voice library (licensed voices designed to work straight out of the box)
- Voice Cloning via short audio clips (5-60 seconds), and the model matches the voice instantly
- Natural and expressive voice synthesis with realistic intonation, rhythm, and emotional range.
- High-fidelity, high-clarity voice output.
- Multilingual support across 23 languages.
- One voice across all supported languages: a single voice identity is rendered consistently in every supported language.
- Fine-grained control of tone and delivery: shape emotion, tone, and delivery at the turn/sentence level, including different speaking styles via SSML.
- State-of-the-art instant Voice Cloning: provide a short audio clip (5-60 seconds) and the model clones the voice instantly, with no fine-tuning required. Prompt quality significantly impacts output, with best results from natural, conversational delivery and moderate energy levels. Access requires Microsoft approval and guardrails are in place to avoid misuse.
- Long-form content generation with improved consistency via chunking and context carryover.
- Multi-speaker support for dialogue and conversational scenarios.
Quick facts
PublisherMicrosoft
AuthorMicrosoft AI
TypeText to speech, Audio generation
LifecyclePreview
Input typetext
Output typeaudio
PricingView pricing