← Engineering Dispatches / Applied AI
Streaming Voice AI: Architecting Sub-300ms Speech-to-Speech WebRTC Pipelines
By Hammad Haider · 11 min read read
Architectural Takeaways
- Eliminate HTTP chunk uploads by streaming raw PCM audio directly over WebRTC media tracks.
- Deploy client-side Voice Activity Detection (Silero VAD) to interrupt AI speech instantly when the human begins speaking.
- Begin text-to-speech audio generation on the very first 4 LLM completion tokens rather than waiting for complete sentence completions.
1. The 300ms Latency Budget: Deconstructing the Pipeline
To feel conversational, total round-trip delay must stay under 350ms. We budget: WebRTC network transit (25ms), Streaming VAD & STT (80ms), First-token LLM generation (100ms), and Streaming Audio TTS buffer (95ms).
2. WebRTC Media Server Architecture with Node.js & Mediasoup
WebRTC sends continuous Opus-encoded audio packets over UDP, bypassing the TCP handshake and head-of-line blocking that plague WebSocket implementations under high packet loss.
3. Token-Pipelined Text-to-Speech Synthesis
Rather than buffering complete sentences, the server streams the initial 5-8 tokens directly to low-latency neural TTS APIs (Cartesia or ElevenLabs Turbo v2.5), playing audio to the user while downstream tokens are still generating.
4. Voice Activity Detection & Zero-Latency Audio Barge-In
When Silero VAD detects user vocalization while the bot is speaking, the client immediately drops the incoming audio buffer and transmits an abort signal over WebRTC DataChannel to cancel active LLM generation.
Read more technical guides on our Dispatches Index →