To support low-latency real-time voice interactions at scale, OpenAI rearchitected its WebRTC infrastructure by decoupling media termination from connection state. By implementing a relay plus transceiver model, the team achieved deterministic first-packet routing and global low-latency connectivity while maintaining standard WebRTC client compatibility.
Key points
Decoupling media relays from session-state-heavy transceivers allows for more flexible and scalable infrastructure deployment.
Using ICE ufrag fields provides an efficient, protocol-native hook for deterministic routing without requiring expensive lookup dependencies.
Optimizing network stacks in Go with kernel-level flags like SO_REUSEPORT and thread pinning can often negate the need for complex kernel bypass solutions.
Preserving standard protocol semantics at the edge ensures interoperability while moving complexity into a thin, centralized routing layer.