Skip to main content
Microsoft Foundry
Azure-Speech-Speech-to-text

Azure-Speech-Speech-to-text

Transcribes streaming or recorded audio into readable text across 140+ languages and dialects. Accuracy can be further optimized with custom models for your specialized use cases.
Microsoft
Version: 1

Azure Speech is a comprehensive suite of AI-powered speech capabilities that includes speech to text, text to speech, speech translation, and voice live AI. It enables developers to build intelligent voice-enabled applications with high accuracy, multilingual support, and customizable voice experiences.

About this model

Speech to text offers various options to transcribe audio data into text.

Real-time speech to text: Instant transcription with intermediate results for streaming audio inputs.

Fast transcription: Fastest synchronous file-based processing for situations with predictable latency.

Batch transcription: Efficient processing for large volumes of prerecorded audio files.

LLM speech (preview): Transcribe and translate audio files using LLM-enhanced speech models, with improved quality and support for prompt tuning.

Custom speech: Fine-tune models with enhanced accuracy for specific domains and use casess.

Key model capabilities

  • Real time streaming, batch, or fast transcription of audio data
  • LLM-powered audio file transcription and translation (preview)
  • Multilingual audio processing
  • Diarization
  • Language identification
  • Word timing
  • Fine tuning

Quick facts

Model providerMicrosoft
TypeAutomatic speech recognition, Speech to text
LifecycleGenerally available (GA)
Input typeaudio
Output typetext