In everyday words
Instead of recording audio and sending it later, an app can talk to the API continuously and get immediate speech-to-text, translation, and spoken replies.
Need a meaning?
An API pattern where audio is streamed continuously so models can respond while a conversation is happening.Speech-to-text that outputs partial text as someone speaks, reducing perceived latency.
Quick Sip
What you need to know
- Who is affected
- developers, operators, product teams
- What changed
- On May 7, 2026, OpenAI introduced three Realtime API audio models: GPT‑Realtime‑2 for voice interactions with stronger reasoning, GPT‑Realtime‑Translate for live speech translation, and GPT‑Realtime‑Whisper for streaming transcription.
- Why it matters
- Realtime voice apps often fail on long context, tool calls, or multilingual use. These models target lower-latency voice agents that can keep a conversation going while translating or transcribing, which can expand where voice interfaces are practical.
- What to watch next
- Watch developer adoption in the Realtime API, how translation quality holds up in noisy settings, and whether voice agents reliably handle interruptions and tool calls in production.
Four useful details
- GPT‑Realtime‑2 targets live conversations with stronger reasoning and longer context for agent workflows.
- GPT‑Realtime‑Translate supports live speech translation across 70+ input languages into 13 output languages.
- GPT‑Realtime‑Whisper provides low-latency streaming transcription priced per minute.
OpenAI · Official AnnouncementOffering Zero Data Retention for frontier models ↗
Adds source-backed context on ai safety from OpenAI.
OpenAI · Official UpdateAdvancing content provenance for a safer, more transparent AI ecosystem ↗Adds source-backed context on ai news from OpenAI.
Notion · Official AnnouncementIntroducing Notion’s Developer Platform ↗Adds source-backed context on ai tools from Notion.
Your next sip
All latest briefings →