Voice Gateway Architecture
How the Blazil Rust media plane handles audio — and how that relates to the latency a caller actually hears.
What our gateway adds while relaying each 20ms audio frame. This is an engineering property, not a promise about response time.
You stop talking → the agent starts speaking. Measured across 400+ real calls; roughly half land under a second.
These measure different things, and the small one does not imply the big one. The 50ms figure is the gateway’s own processing cost per frame; the 1.2–1.5s figure is the end-to-end conversational turn. Anyone quoting a millisecond number as a voice agent’s response time is describing their plumbing, not their product.
Why this matters at all: audio arrives every 20ms and must leave on the same cadence. A gateway that cannot keep up buffers, and buffering is the 800ms+ of dead air that makes traditional cloud telephony feel like a walkie-talkie. Staying under the frame budget is what buys the headroom — it is a floor beneath the numbers on the right, not a substitute for them.
Knowing you finished. A predictive endpointing model decides you have stopped rather than paused mid-sentence. Cut this too fine and the agent talks over people — the single most unnatural thing a voice agent can do.
Thinking. Routing to the right agent, retrieving from your knowledge base, and generating the reply.
Speaking. Time to the first audio from the voice model. Generation streams, so the agent begins talking before the sentence is finished.
Barge-in runs throughout: interrupt at any point and the agent stops immediately, which in practice matters more to how fast a conversation feels than any of the three numbers above.