OpenAI gpt-realtime-whisper
OpenAI gpt-realtime-whisper
Version: 2026-05-07
OpenAILast updated May 2026
A new STT (speech to text) model with realtime capability.

Direct from Azure models

Direct from Azure models are a select portfolio curated for their market-differentiated capabilities:
  • Secure and managed by Microsoft: Purchase and manage models directly through Azure with a single license, consistent support, and no third-party dependencies, backed by Azure's enterprise-grade infrastructure.
  • Streamlined operations: Benefit from unified billing, governance, and seamless PTU portability across models hosted on Azure - all part of Microsoft Foundry.
  • Future-ready flexibility: Access the latest models as they become available, and easily test, deploy, or switch between them within Microsoft Foundry; reducing integration effort.
  • Cost control and optimization: Scale on demand with pay-as-you-go flexibility or reserve PTUs for predictable performance and savings.
Learn more about Direct from Azure models .

Key capabilities

About this model

Gpt-realtime-whisper is a low‑latency, streaming speech‑to‑text model designed for real‑time transcription of live audio. It continuously processes incoming audio streams and converts spoken language into text with high accuracy, making it well suited for conversational AI, voice assistants, and live captioning scenarios. The model is optimized for robustness across diverse accents, speaking styles, and acoustic conditions, enabling reliable transcription in dynamic, real‑world environments while maintaining minimal latency for interactive applications. Supported region: Canada Central, France Central, and India South. More coming soon

Key model capabilities

Key Features:
  • Real-time speech-to-text transcription
    Continuously converts streaming audio into text during live interactions.
  • Low-latency streaming performance
    Processes incoming audio incrementally to support interactive and near real-time applications.
  • Optimized for live audio input
    Designed to handle continuous microphone or call audio streams rather than batch uploads.
    -Robust speech recognition
    Transcribes spoken language across varied speaking styles and environments.
  • Supports conversational pipelines
    Enables downstream use cases such as voice agents, live captioning, and transcription-driven workflows.
  • Text-only output (no audio generation)
    Produces structured transcription output without generating synthesized speech.
  • Integrates with realtime model stack
    Can be paired with translation or conversational models to build full end-to-end voice experiences.

Use cases

See Responsible AI for additional considerations for responsible use.

Key use cases

The provider has not supplied this information.

Out of scope use cases

The provider has not supplied this information.

Pricing

Pricing is based on a number of factors, including deployment type and tokens used. See pricing details here.

Technical specs

The provider has not supplied this information.

Training cut-off date

The provider has not supplied this information.

Training time

The provider has not supplied this information.

Input formats

The provider has not supplied this information.

Output formats

The provider has not supplied this information.

Supported languages

The provider has not supplied this information.

Sample JSON response

The provider has not supplied this information.

Model architecture

The provider has not supplied this information.

Long context

The provider has not supplied this information.

Optimizing model performance

The provider has not supplied this information.

Additional assets

The provider has not supplied this information.

Training disclosure

Training, testing and validation

The provider has not supplied this information.

Distribution

Distribution channels

This model is provided through the Azure OpenAI Service.

More information

The following documents are applicable:

Responsible AI considerations

Safety techniques

The provider has not supplied this information.

Safety evaluations

The provider has not supplied this information.

Known limitations

The provider has not supplied this information.

Acceptable use

Acceptable use policy

The provider has not supplied this information.

Quality and performance evaluations

Source: OpenAI The provider has not supplied this information.

Benchmarking methodology

Source: OpenAI The provider has not supplied this information.

Public data summary

Source: OpenAI The provider has not supplied this information.
Model Specifications
Context Length128000
LicenseCustom
Training DataApril 2026
Last UpdatedMay 2026
Input TypeAudio
Output TypeText
ProviderOpenAI
Languages27 Languages