gpt-realtime-translate

Gpt‑realtime‑translate is a low‑latency streaming model that converts spoken audio into translated output in real time, enabling live cross‑language communication within voice applications.

Azure OpenAI

Direct from Azure

Version: 2026-05-07

Direct from Azure models

Direct from Azure models are a select portfolio curated for their market-differentiated capabilities:

Secure and managed by Microsoft: Purchase and manage models directly through Azure with a single license, consistent support, and no third-party dependencies, backed by Azure's enterprise-grade infrastructure.
Streamlined operations: Benefit from unified billing, governance, and seamless PTU portability across models hosted on Azure - all part of Microsoft Foundry.
Future-ready flexibility: Access the latest models as they become available, and easily test, deploy, or switch between them within Microsoft Foundry; reducing integration effort.
Cost control and optimization: Scale on demand with pay-as-you-go flexibility or reserve PTUs for predictable performance and savings.

Learn more about Direct from Azure models .

Key capabilities

About this model

Gpt-realtime-translate is a low‑latency, streaming model for real‑time speech translation, designed to convert spoken language into translated output during live audio interactions. It processes continuous audio streams and enables applications to translate speech across languages as it is spoken, supporting multilingual communication scenarios such as live conversations, voice assistants, and cross‑language interactions. The model is part of a broader set of speech capabilities that include transcription and translation, allowing developers to build end‑to‑end voice pipelines that operate in real time.

Supported region: Canada Central, France Central, and India South. More coming soon

Key model capabilities

Key Features:

Real-time speech-to-speech translation
Converts incoming audio into translated speech output during live, streaming interactions.
Simultaneous translated transcription
Always provides a text transcript of the translated audio in the target language.
Low-latency streaming operation
Processes continuous audio input and returns translated audio in small streaming chunks for near real-time responsiveness.
Continuous audio input handling
Designed for ongoing audio streams, including pauses (expects continuous input rather than discrete clips).
Target-language translation control
Translates speech into a specified output language configured per session.
Optional input transcription (via separate model)
Can include source-language transcripts, enabled through an external transcription model integrated into the pipeline.
Streaming audio + text outputs
Emits both audio and transcript deltas incrementally, enabling synchronized playback and display.

Use cases

Pricing

Technical specs

Training disclosure

Distribution

More information

Quick facts

Model providerAzure OpenAI

TypeSpeech translation

LifecycleGenerally available (GA)

Input typeaudio

Output typeaudio, text

Context window128k

Token limits4096 output

PricingView pricing

gpt-realtime-translate

About this model

Key model capabilities

Quick facts

Quick start