MarkTechPost
7/7/2026

OpenAI Releases GPT-Realtime-2.1 and GPT-Realtime-2.1-mini for Low-Latency Voice Agents in the API
Short summary
OpenAI released GPT-Realtime-2.1 and GPT-Realtime-2.1-mini models for building low-latency voice agents through the OpenAI API. Both models achieve at least 25% lower p95 latency improvements through enhanced caching strategies. The mini variant maintains pricing parity with the earlier gpt-realtime-mini, while WebRTC integration enables seamless real-time voice connections for developers building voice-first applications.
- •Two new Realtime models released for voice agents
- •25% latency improvement via enhanced caching
- •Mini model priced identically to predecessor
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



