OpenAI voice intelligence features are getting a major upgrade as the company expands its API with new tools focused on real-time conversations, transcription, and live translation. On Thursday, OpenAI announced several additions designed to help developers build apps that can talk with users, translate conversations instantly, and convert speech into text as discussions happen.
The launch introduces three main capabilities: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper. According to OpenAI, the new models are intended to move voice interfaces beyond simple responses and toward systems that can actively understand, reason, and take action during live conversations.
OpenAI Voice Intelligence Features Add Smarter Real-Time Conversations
One of the biggest additions is GPT-Realtime-2, OpenAI’s latest conversational voice model.
The company said the model is designed to create more natural vocal interactions while also handling more advanced user requests. Unlike the earlier GPT-Realtime-1.5 model, GPT-Realtime-2 is built with GPT-5-class reasoning capabilities, which OpenAI says allows it to better manage complex conversations and multi-step interactions.
The goal appears to be making voice-based AI assistants feel more responsive and capable during live discussions instead of operating like traditional call-and-response systems.
Real-Time Translation Now Supports More Than 70 Languages
Another major part of the OpenAI voice intelligence features rollout is GPT-Realtime-Translate.
The feature is designed to deliver live conversational translation that can keep pace with ongoing discussions. OpenAI said the model currently supports more than 70 input languages and 13 output languages.
That means the system can understand a large range of spoken languages while translating responses into supported target languages in real time.
The company positioned the feature as a tool that could support a variety of use cases where multilingual communication is important.
GPT-Realtime-Whisper Brings Live Speech-to-Text Capabilities
OpenAI also introduced GPT-Realtime-Whisper, a new speech-to-text feature built for live transcription.
According to the company, the model captures spoken conversations and converts them into text as interactions happen in real time.
The feature expands OpenAI’s growing focus on audio-based AI experiences and could help developers create tools centered around accessibility, communication, and workflow automation.
OpenAI said the combined release of these models moves real-time audio systems closer to interfaces that can “listen, reason, translate, transcribe, and take action as a conversation unfolds.”
OpenAI Sees Enterprise and Creator Use Cases
The company said the new tools could be useful across multiple industries and platforms.
Customer service appears to be one of the clearest targets, especially for companies looking to improve voice-based support systems. However, OpenAI also said the new capabilities may support use cases in education, media, events, and creator-focused platforms.
Because the tools are integrated into the Realtime API, developers can build conversational audio experiences directly into their own applications and services.
Safety Guardrails Included in New Audio Models
OpenAI acknowledged that advanced voice systems could potentially be misused for spam, fraud, or other harmful online activity.
According to the company, the new OpenAI voice intelligence features include built-in safeguards intended to reduce abuse. OpenAI said the systems contain triggers that can halt conversations if interactions are detected as violating the company’s harmful content guidelines.
The company also explained that GPT-Realtime-Translate and GPT-Realtime-Whisper will be billed by the minute, while GPT-Realtime-2 pricing will be based on token usage.
Don’t miss out on our latest news—follow us for the latest AI news, breakthroughs, and insights that matter.
