Azure-Speech-Speech-to-text
Azure Speech is a comprehensive suite of AI-powered speech capabilities that includes speech to text, text to speech, speech translation, and voice live AI. It enables developers to build intelligent voice-enabled applications with high accuracy, multilingual support, and customizable voice experiences.
About this model
Speech to text offers various options to transcribe audio data into text.
Real-time speech to text: Instant transcription with intermediate results for streaming audio inputs.
Fast transcription: Fastest synchronous file-based processing for situations with predictable latency.
Batch transcription: Efficient processing for large volumes of prerecorded audio files.
LLM speech (preview): Transcribe and translate audio files using LLM-enhanced speech models, with improved quality and support for prompt tuning.
Custom speech: Fine-tune models with enhanced accuracy for specific domains and use casess.
Key model capabilities
- Real time streaming, batch, or fast transcription of audio data
- LLM-powered audio file transcription and translation (preview)
- Multilingual audio processing
- Diarization
- Language identification
- Word timing
- Fine tuning