Skip to main content
Omazy Engineering

RFC 2232 · voice latency

From a silent call to 820 milliseconds.

A bilingual voice agent went live and callers hung up. Some calls connected and played nothing at all; the ones that worked were too slow to feel like a conversation. This is the measured path from there to a greeting that starts in 820 milliseconds, including the two experiments that did not pay off and the crash we shipped ourselves.

// measured on live calls · Bengali and English // every number here came off a live call
820 ms
greeting time to first audio, down from 2.9 s
59%
of the caller wait was speech synthesis
12.2 ms
of synthesis per character of reply
0.6 s
best case from swapping the language model
429 KB
the entire fixed-phrase cache, 21 phrases

Section 01 · the result first

The greeting was the easiest 2 seconds on the whole call.

Every call opens with the same sentence, in the same voice, in the same language. It was being synthesized from scratch every single time, at 2.2 to 2.7 seconds of vendor round trip, for audio we had already paid for on the previous call.

Greeting · time to first audio

before 2.90 s
after 0.82 s

The darker inset on the lower bar is the 610 ms WebSocket handshake, which is a floor no cache can move. Everything to the right of it is what the cache removed.

Getting there took a long stretch of work that was mostly not about latency. Before anything could be timed, calls had to stop being silent, the agent had to remember what the caller said one turn ago, and it had to stop discarding answers to its own questions. Then the profiling could start, and the profiling produced two results that pointed the opposite way from the intuition.

Section 02 · connected and silent

Four separate bugs produced exactly the same symptom.

A call that rings, answers, and then plays nothing is the least informative failure in this stack, because every stage can cause it and none of them says so. We hit four causes in a row, and each one was invisible until the one in front of it was fixed.

A model id that did not exist

The transcriber was configured with a speech-to-text model name that the vendor does not publish. The session opened, accepted audio, and returned nothing forever. Nothing in the transport layer treats "no transcripts yet" as an error, because on a quiet line that is the correct state.

A voice that refused the language

The synthesis model chosen for the deployment does not accept Bengali. It rejected each request individually. The rejection was real, and the error carrying it was being discarded one layer up, which is the next bug.

No greeting at all

The agent waited for the caller to speak first. That means the caller hears the entire recognition, model and synthesis chain as dead air before anything happens. Callers hung up at 13 seconds, which is generous.

The stream was never opened

The root cause. The outbound media stream has to be opened with a start event before any audio frame is accepted. We were sending media frames straight down a socket that was in the right state to receive them and had never been told a stream existed, so the platform discarded every one, silently, with no error frame and no close.

The line that made all four hard to see

Three call sites in the turn loop finished a synthesis with a discarded error. That reads as harmless, and on a streaming synthesizer it nearly is. On a request and response synthesizer it is the opposite: the terminal call is where the vendor request actually happens, so it is the only place a rejection can surface.

// Done() is where a request/response synthesizer actually calls the vendor, so
// this is where a TTS failure surfaces. Discarding it meant a rejected
// synthesis produced zero audio, no error and no log: the runner answered
// 200, the turn "succeeded", and the caller heard nothing.
var ttsErr error
if opened {
    ttsErr = sink.tts.Done()
    <-drained // all audio enqueued
}

With the error dropped, a rejected synthesis produced zero audio, no error, and a turn that reported success. The turn counter went up. The latency metric recorded a healthy number. The caller heard nothing. Any failure path that increments a success metric will outlive several rounds of debugging, because it removes itself from the list of suspects.

The general rule we took from this: in a media pipeline, treat "produced no output" as an error condition with its own alarm, separately from "returned an error". The two are the same event to a caller and completely different events to the code.

Section 03 · state

Every caller utterance was turn one of a brand new conversation.

The endpoint that runs a turn minted a fresh session per request. Each utterance therefore arrived at the language model with no history, as the opening line of a conversation that had never happened.

This did not arrive as one bug report. It arrived as three, from three different people, that read like three unrelated quality problems:

One cause. The fix is that the bridge, which is the only component that lives for the length of the call, keeps the transcript and sends it with every turn. That is the right place for it: the turn endpoint is stateless by design and the media socket is the natural lifetime of a call.

Worth naming, because it recurs: three convincing, specific, independently reported symptoms can share one line of cause. Fixing them as three tickets would have produced three workarounds and left the session bug in place.

Section 04 · turn-taking

The agent asked a question and threw the answer away.

A final transcript that landed while the agent was still speaking failed a state-machine guard, and the utterance was discarded with a single line in the log. On a phone call that is the caller answering a direct question and being ignored.

sequenceDiagram
  participant C as Caller
  participant S as Session state
  participant A as Agent
  Note over C,A: before · the guard wins
  C->>S: final transcript arrives while agent is Speaking
  S-->>A: BeginTurn refused
  A->>A: utterance dropped, one line in the log
  Note over C,A: after · the caller wins
  C->>S: final transcript arrives while agent is Speaking
  A->>S: bargeIn · Speaking becomes Listening
  A->>S: BeginTurn accepted
  A->>C: answers what the caller actually said
Fig 1. The guard was correct about the state and wrong about the priority. A transcript is proof that the caller spoke, so the agent yields the floor first and takes the turn second.

Overlap is normal on a phone call, not an error. The detector at the telephony edge usually clears playback before this happens, but it only fires while the platform reports the agent stream as playing, and the pacer on the bridge can still be draining. There is a window on either side of every media start and stop event where the two disagree about who is talking.

The corrected order is: cut the audio, move the session out of speaking, then begin the turn. One state that genuinely cannot overlap with itself is thinking, because final transcripts are handled serially, but a greeting or an idle prompt synthesizes from its own goroutine. A short bounded wait covers that race instead of dropping what the caller said.

Section 05 · the one we caused

A nil pointer, a full green test suite, and every call panicking twice.

The clock is injected so that timing can be tested without sleeping. It was defaulted inside a constructor, and the constructor took its configuration struct by value.

// withDefaults fills in what an injected Deps may legitimately leave unset.
//
// It exists because the clock used to be defaulted inside newAgent, on
// newAgent's own by-value COPY of Deps. That left the caller's copy with a
// nil clock, and RunAgent, which also calls deps.clock(), dereferenced nil
// and took the whole call down with it. Normalize once, at every entry
// point, before anything reads a field.
func (d Deps) withDefaults() Deps {
    if d.clock == nil {
        d.clock = time.Now
    }
    return d
}

The constructor filled in its own private copy. The caller kept the copy with the nil field, and the request handler, which calls the clock directly before the constructor ever runs, dereferenced it. Every call panicked twice, the socket died, and the telephony scenario did the only sensible thing it could: it fell back to the recorded apology about a technical problem and asked the caller to try later.

The interesting part is why the tests passed

Coverage on this package was good. Every unit test constructed the inner object directly and injected a fake clock, which is exactly what you want for testing a pacer against fixed timestamps. Not one of them went through the real request handler, because the real handler takes a concrete WebSocket type and standing one up in a test is more work than the test being written at the time was worth.

So the tests proved the object behaved correctly when constructed the way tests construct it. The panic lived entirely in the gap between that and the way production constructs it, which no test occupied.

The fix in the code

Normalise the configuration at every entry point rather than in one constructor, and do it before any field is read. A defaulting step that only one of two entry points calls is not a default, it is a coin flip.

The fix in the tests

Stand the real handler behind a listener and drive it over a real socket, with no injected clock, asserting that a call reaches a greeting. Awkward to write once, and it covers the assembly seam that unit tests structurally cannot reach.

If a component is only ever built by its tests one way and by production another way, the tests are describing a different program. That is worth a targeted test at the seam, and it does not need to be a full end-to-end suite to be worth having.

Section 06 · endpointing

The defaults chopped one Bengali sentence into three.

Reported as "the voice processing is jittery". The recording showed one continuous utterance registering as three: speech start at 101.4, end at 103.8, start again at 104.3, end at 105.8.

Every fragment ends a turn. So the agent answered the first third of a sentence, started speaking, and then barged in on itself when the second third arrived. From the caller side that is not a tuning problem, it is an agent that cannot hold a thought.

// Tuned against a real bilingual call where the defaults chopped one
// utterance into three:
//   speechStart 101.4 / end 103.8 / start 104.3 / end 105.8
// Every fragment ends a turn, so the agent answers half a sentence and then
// barges in on itself. That is most of what "the voice processing is
// jittery" actually was.
vad: { threshold: 0.6, minSilenceDurationMs: 700, speechPadMs: 150 },
turnDetector: { threshold: 0.5 },
bargeInGuardMs: 400,
KnobWasNowWhy
threshold0.50.6At 0.5 the line noise and the room tone register as speech, so the detector barges in on the agent for nothing.
minimum silence300 ms700 msBengali carries longer intra-sentence pauses than English, and a caller reading a phone number pauses between groups of digits. At 300 ms the turn ends mid-number.
speech padding10 ms150 ms10 ms clips the onset of the first word, which is the syllable the recognizer needs most.
barge-in guardnone400 msWithout it the tail of speech from the caller, and the agent voice coming back through a speakerphone, cuts the reply before its first word is out.

The language matters here and the defaults do not know it. Bengali carries longer pauses inside a sentence than English does, and any caller reading out digits pauses between groups regardless of language. A 300 millisecond silence window ends the turn in the middle of a phone number, which is precisely the moment in a booking flow where you can least afford it.

These four values are per-deployment configuration, not constants. A tuning that is correct for one language and one line quality is a guess for the next one, so they live in the call scenario config with the reasoning written next to them.

Section 07 · geography

Identical build, two calls minutes apart, opposite outcomes.

Same binary, same configuration, same agent, same phone number. The only difference was the network the caller dialled from, and it decided which region the telephony platform allocated a media server in.

MeasurementCall A · workedCall B · failed
caller networklocal ISPVPN
media server regionGermanySingapore
bridge regionus-eastus-east
WebSocket handshake0.48 s3.1 s, then 8.3 s
speech events100
agent audio6 repliesunder 1 second in total
outcome93 second conversationabandoned at 15 seconds

Call B routed audio from a Singapore media server to a bridge in us-east and back. The handshake alone took 3.1 seconds and then 8.3 seconds on the retry. Not one speech event was ever detected, the agent managed under a second of audio in total, and the caller gave up at 15 seconds. There was no bug to find, and we spent real time looking for one.

Check the media server region before you touch the code. A voice bug that reproduces on one network and not another is a topology question until proven otherwise. This is also the argument for co-locating the bridge with the telephony point of presence: it is the only part of the budget that is pure geometry.

Section 08 · the measurement

Synthesis is 59 percent of the wait, and it is linear in characters.

Once calls were reliable, five representative questions were run against the live stack and each stage timed. Mean language model time was 1.45 s. Mean synthesis time was 2.07 s. Synthesis is the larger half of everything a caller waits through.

flowchart LR
  C["Caller stops speaking"] --> V["Turn detector
700ms of silence"] V --> S["Speech to text
final transcript"] S --> L["Language model
mean 1.45s"] L --> T["Speech synthesis
mean 2.07s"] T --> P["Pacer
one 20ms frame per 20ms"] P --> A["Caller hears the reply"]
Fig 2. One turn, with the two measured stages marked. Everything before the model is either fixed by the detector configuration or bounded by the network; everything after it scales with how much the agent decided to say.

Per question · model time and synthesis time

language model speech synthesis
where are you located 2.57 s

model 1.32 s · synthesis 1.25 s · reply 83 characters

what are your opening hours 2.07 s

model 0.99 s · synthesis 1.08 s · reply 69 characters

I would like an appointment 1.67 s

model 1.20 s · synthesis 0.47 s · reply 19 characters

what should I bring with me 5.91 s

model 1.80 s · synthesis 4.11 s · reply 317 characters

can you help with something from before 5.37 s

model 1.92 s · synthesis 3.45 s · reply 263 characters

The useful answers are the slow ones

Synthesis time on this vendor fits a straight line: 0.24 seconds of fixed cost plus 12.2 milliseconds per character. That fit reproduces all five measurements to the nearest hundredth of a second, so it is a good enough model to plan against.

Caller questionModelReplySynthesisTotal
where are you located 1.32 s 83 c 1.25 s 2.57 s
what are your opening hours 0.99 s 69 c 1.08 s 2.07 s
I would like an appointment 1.20 s 19 c 0.47 s 1.67 s
what should I bring with me 1.80 s 317 c 4.11 s 5.91 s
can you help with something from before 1.92 s 263 c 3.45 s 5.37 s

The consequence is uncomfortable. A 19 character reply costs 0.47 s of synthesis; a 317 character reply, which is the one that actually answers the question about what to bring, costs 4.11 s. The most useful answers are systematically the slowest, and the slowness is proportional to how useful they are. Any optimisation that works by shortening replies is trading the answer for the timing.

Section 09 · negative result one

Changing the language model buys 0.6 seconds, mostly by saying less.

The instinct when a turn feels slow is to reach for a faster model. Five were run against the same five questions on the same stack.

Same five questions · mean total time per turn

gpt-4o-mini 4.03 s
gpt-4.1-mini 4.57 s
gpt-4o 4.66 s
gpt-4.1 4.68 s
claude-haiku-4.5 6.40 s

The deployed default sits at 4.66 s. The fastest option in the set is 0.63 s quicker, and the slowest is 1.74 s behind the default.

A 0.6 second improvement is real, and it is not where the problem is. More to the point, the spread across these five is not a spread in raw generation speed. It tracks how long each model chose to make its replies, and reply length is the input to the stage that dominates the budget. The model choice is mostly a proxy for verbosity.

This reframes the question. Instead of "which model is fastest", the question is "what determines reply length, and can we control it without controlling the answer". Which leads directly to the next experiment, and to the wrong answer.

Section 10 · negative result two

We told the model to be brief and it got slower.

If synthesis is linear in characters, capping characters caps synthesis. So the agent brief gained an explicit instruction: at most two sentences, at most 200 characters. Then the same five questions were run again.

With an explicit two-sentence, 200-character cap in the prompt

list answer · what to bring+46 characters

no cap 317 c
capped 363 c

list answer · multi-step question+12 characters

no cap 263 c
capped 275 c

mean synthesis time, all five+0.27 s

no cap 2.07 s
capped 2.34 s

Each pair is scaled to its own larger value. The first two rows are reply length in characters, the third is mean synthesis time across all five questions.

The short answers stayed short, in the 69 to 84 character band they were already in, because they were never the problem. The two list-shaped answers, the ones the cap was written for, both got longer. Mean synthesis time went up from 2.07 s to 2.34 s. The change was reverted.

We did not chase the mechanism far enough to be certain of it, and the honest summary is the measurement: an instruction of this kind is not a reliable control surface for output length on a list-shaped question. If reply length has to be bounded, it needs to be bounded by something deterministic outside the model, or the answer needs to be restructured so brevity is natural rather than requested.

Recording this mattered more than the 0.27 s it cost. Left unmeasured, "we tell it to be concise" would have stayed in the brief as a latency control that everyone believed in and nothing supported.

Section 11 · the win

Some of the audio is byte-identical on every call, so pay for it once.

The greeting, the acknowledgements, the idle nudges, the farewell, the line that asks the caller to repeat themselves. Every one is a fixed string in a fixed voice, and every one was being re-synthesized on every call.

flowchart TB
  T["Synthesize some text"] --> D{"marked cacheable?"}
  D -->|"no · a generated reply"| V["vendor round trip"]
  D -->|"yes · a fixed phrase"| K["key = text + language + voice"]
  K --> H{"already rendered?"}
  H -->|"hit"| OUT["audio, no network"]
  H -->|"miss"| V
  V --> ST["store, if cacheable"]
  ST --> OUT
  BOOT["boot · background warm"] -.->|"pre-render 21 phrases"| K
Fig 3. A caching synthesizer wrapping the vendor driver. The key is text plus language plus voice, because the same words in a different voice are different audio, and serving one caller the voice another caller chose is exactly the bug this must not have.

Caching everything is the trap

The obvious version of this feature caches all synthesis output. It is worse than useless. Generated replies are never byte-identical twice, so an everything-cache pays full storage for a hit rate near zero, and then evicts the handful of fixed phrases that would have hit. So caching is opt-in per phrase: the caller marks a request cacheable, and nothing else is stored.

That also removes the need for an eviction policy. The cacheable set is a small fixed list known at configuration time, so a least-recently-used policy could only ever discard something about to be needed again. There is a hard entry cap as a guard against someone marking something unbounded as cacheable, and past it synthesis still works, it just stops being stored.

Storage is never the constraint

21 phrases across 3 voices is 429 KB in total. A 13 second Bengali greeting is about 104 KB of 8 kHz mu-law. The cache ceiling is 8 MiB, which is more fixed phrases than any agent has, and the whole thing fits in the memory a single call already uses.

Warm it at boot, in the background

Pre-rendering the phrase set at startup means the first caller of the day does not pay for it. Doing that warm before starting the listener left the container up and refusing sockets for 15 seconds, which the telephony scenario correctly turns into the technical-problem apology. Warm on a goroutine and serve immediately.

Instant acknowledgements

The second half is cheaper and does more for how the call feels. A turn costs about 1.5 s of model plus 2 s of synthesis before the caller hears a syllable, and up to 6 s on the long answers. That silence is what makes a caller repeat themselves or start talking over the reply.

sequenceDiagram
  participant C as Caller
  participant B as Bridge
  participant L as Language model
  participant T as Synthesis
  C->>B: stops speaking
  B->>C: acknowledgement from cache, about 1s
  par behind the acknowledgement
    B->>L: run the turn
    L-->>B: reply text
    B->>T: synthesize the reply
    T-->>B: audio
  end
  B->>C: reply plays straight on the end of it
Fig 4. A one second acknowledgement, played from cache the moment the caller stops, queued ahead of the reply inside the same turn. The reply lands behind it and the two play as one continuous utterance.

This does not make the answer arrive sooner. It makes the caller stop waiting in silence, which is the part they experience. A human receptionist does exactly this. Two properties are load-bearing: the phrases rotate, because the same acknowledgement before every answer is a sharper robot tell than the silence it replaced, and they stay short, because the acknowledgement plays before the reply so its length is added to the total.

It is queued straight into the pacer rather than going through the normal speak path, because that path opens its own turn and the acknowledgement has to live inside the turn that is already running. Failures are swallowed on purpose: comfort noise must never fail the turn that carries the actual answer.

Section 12 · what is next

Streaming synthesis, because it helps every turn.

The synthesis driver currently uses a request and response endpoint. It sends the whole reply, waits, and gets the whole audio back. A 363 character reply therefore waits the full 4.7 seconds before one byte of it plays.

The same vendor offers a streaming WebSocket. Audio would start after the first sentence rather than after the last one. On the linear fit that is roughly 3.5 seconds off the long answers and 1.5 to 2 seconds off the average. Unlike the cache, which only helps phrases that repeat, this helps every generated reply, which is every turn that matters.

What the cache cannot do

It removed a 2 second fixed cost from the opening of every call and it made the acknowledgements free. It cannot touch the answer itself, because the answer is different every time. The remaining budget is almost entirely generated audio.

The floor underneath all of it

610 ms of the 820 ms greeting is the WebSocket handshake. That is the shape of the remaining problem: past a point, the wins stop being software and start being geography, which is Section 07 again.

Related reading

The architecture underneath all of this, the media bridge, the pacer, the barge-in state machine, the bilingual cascade and the interface that lets the phone agent reuse the chat engine, is written up in Voice AI. The same driver-and-factory shape the synthesis vendors sit behind is described for messaging in the Common Channel Wrapper, and the per-turn model, latency and cost numbers operators see come from Conversation Observability.

// RFC 2232 · voice latency · numbers from live calls

// edit, open a PR, ships on merge to main