Triamorph Systems

← Engineering Dispatches / Applied AI

Streaming Voice AI: Architecting Sub-300ms Speech-to-Speech WebRTC Pipelines

By Hammad Haider · 11 min read read

Human conversation requires response latencies under 400 milliseconds. Traditional sequential voice setups (record audio -> upload MP3 -> transcribe -> send prompt -> wait for full completion -> synthesize audio) routinely suffer 3 to 5 second pauses, breaking conversational immersion. This blueprint details how to achieve sub-300ms end-to-end voice latency using WebRTC tracks and streaming pipeline pipelining.

Architectural Takeaways

  • Eliminate HTTP chunk uploads by streaming raw PCM audio directly over WebRTC media tracks.
  • Deploy client-side Voice Activity Detection (Silero VAD) to interrupt AI speech instantly when the human begins speaking.
  • Begin text-to-speech audio generation on the very first 4 LLM completion tokens rather than waiting for complete sentence completions.

1. The 300ms Latency Budget: Deconstructing the Pipeline

To feel conversational, total round-trip delay must stay under 350ms. We budget: WebRTC network transit (25ms), Streaming VAD & STT (80ms), First-token LLM generation (100ms), and Streaming Audio TTS buffer (95ms).

2. WebRTC Media Server Architecture with Node.js & Mediasoup

WebRTC sends continuous Opus-encoded audio packets over UDP, bypassing the TCP handshake and head-of-line blocking that plague WebSocket implementations under high packet loss.

3. Token-Pipelined Text-to-Speech Synthesis

Rather than buffering complete sentences, the server streams the initial 5-8 tokens directly to low-latency neural TTS APIs (Cartesia or ElevenLabs Turbo v2.5), playing audio to the user while downstream tokens are still generating.

4. Voice Activity Detection & Zero-Latency Audio Barge-In

When Silero VAD detects user vocalization while the bot is speaking, the client immediately drops the incoming audio buffer and transmits an abort signal over WebRTC DataChannel to cancel active LLM generation.

Read more technical guides on our Dispatches Index →