MAI-Transcribe-2
The second generation of our speech-to-text model family: more accurate, faster, and built to handle a wider range of real-world audio. MAI-Transcribe-2 transcribes reliably across accents, speaking styles, and noisy environments, and now covers 60 languages with automatic language detection for multilingual recordings.
About this model
MAI-Transcribe-2 is a speech-to-text model built in-house by the Microsoft AI team, delivering reliable transcription across 60 languages. It powers video captions, meeting transcription, accessibility tools, call analysis, content creation workflows, and voice agents. Beyond raw transcription, this generation adds speaker diarization to separate and label multiple speakers in a single recording, and word-level timestamps that pinpoint exactly who spoke and when. Keyword biasing makes transcription domain-aware, improving recognition of industry and scientific terms, proper names, and other specialized vocabulary.
Key capabilities
- Best-in-class transcription accuracy across 60 languages.
- Speaker diarization: identifies and separates multiple speakers within the same audio recording.
- Word-level timestamps: precise timing for every word, enabling alignment, search, and editing workflows.
- Keyword biasing to handle domain specific terminology, abbreviations or disambiguating terms.
- Configurable transcription style:
verbatimcaptures speech exactly as spoken, including fillers and false starts, for compliance and analysis workloads;cleanremoves disfluencies for readable captions, notes, and published transcripts. - Automatic language identification, including multilingual recordings.
- Robust in noisy, real-world conditions and across diverse accents and dialects.
- Faster inference with substantially lower latency, especially for long‑form audio, with up to 10× faster processing than competitors.