RFC 2232 · Voice AI
Voice AI
How a phone call runs as a real-time turn loop. This article covers the four parts of the system: the media bridge, turn-taking and barge-in, the bilingual speech cascade, and reusing the chat engine behind an importable interface.
Overview
The pipeline
A voice call has a hard timing constraint. The gap between the caller finishing a sentence and the agent starting to reply has to stay short, roughly under a second and a half at the median, or the caller talks over the agent. The system keeps that number low while reusing the same conversation engine as chat.
flowchart LR
subgraph EDGE["Edge · co-located"]
VOX["Telephony POP"]
BR["Media bridge"]
VOX <-->|"mu-law / 8k"| BR
end
subgraph CASCADE["Cascade (off-box)"]
STT["STT"]
LLM["LLM · chat engine"]
TTS["TTS"]
end
subgraph WARM["Platform · warm path"]
DB["chat_sessions
call_records"]
end
BR --> STT --> LLM --> TTS --> BR
BR -.->|"signed control webhook"| DB
Part 01
The media bridge
The bridge is a small, stateless Go service (cmd/voicebridge) that runs close to the telephony point of presence. It terminates the media WebSocket and runs the codec, the jitter buffer, the pacer, and the barge-in state machine. The expensive work runs elsewhere.
flowchart LR
subgraph EDGE["Edge · co-located"]
VOX["Telephony POP
SIP / PSTN + call scenario"]
BR["Media bridge
cmd/voicebridge"]
VOX <-->|"mu-law / 8k over WSS"| BR
end
subgraph BRAINS["Off-box vendors"]
STT["STT"]
LLM["LLM"]
TTS["TTS"]
end
BR --> STT --> LLM --> TTS --> BR
Co-location is a requirement, not an optimisation. The language model is the last stage of the pipeline, so any round trip added there is added to every turn. A bridge running far from the telephony POP spends its budget on network latency before the model produces a token. The bridge runs next to the POP; the durable platform can run anywhere because it is never on the audio path.
Mu-law end to end
Telephony audio is G.711 mu-law at 8 kHz. The bridge keeps that format for the whole round trip: caller audio arrives as mu-law, the transcriber accepts mu-law, the synthesizer is asked to return mu-law, and it goes back to the caller as mu-law. Skipping transcoding removes CPU work and latency. The one remaining per-packet cost is the base64 JSON framing the transport uses for each 20 millisecond packet, and the bridge uses pooled buffers so that framing adds no garbage-collection pressure. The frame protocol, the codec, a jitter-reorder buffer with packet-loss concealment, and the pacer live in one dependency-free package that is unit-tested without a socket.
The turn loop
A call is a sequence of turns. The bridge streams caller audio to speech-to-text. When speech-to-text reports the caller has finished, the bridge runs one turn and streams the reply into text-to-speech, whose audio the pacer sends back one frame at a time.
sequenceDiagram participant C as Caller participant B as Media bridge participant S as STT participant R as Turn runner participant T as TTS C->>B: audio frames (mu-law, 20ms) B->>S: stream audio S-->>B: final transcript B->>R: run one turn (text in) R-->>T: reply tokens (streamed) T-->>B: audio (mu-law) B->>C: paced 20ms frames out
A session state machine tracks the call through listening, thinking, and speaking, and records the per-turn latency breakdown. The pacer owns the transition back to listening, which is covered in Part 02.
Part 02
Turn-taking and barge-in
Turn-taking needs two separate signals, and conflating them is a common mistake. One says the caller has started speaking, so the agent should stop (barge-in). The other says the caller has finished, so it is the agent's turn (endpointing).
Barge-in
A voice-activity detector. When there is speech energy while the agent is talking, stop the agent. It has to be instant, so it runs at the edge, next to the audio.
Endpointing
Deciding the caller is done is a judgement call. A naive silence timer cuts people off mid-thought. A semantic turn detector asks whether the utterance sounds finished, which matters for code-switching speakers who pause while changing language.
Barge-in at the edge
sequenceDiagram participant C as Caller participant E as Edge (VAD) participant B as Media bridge Note over C,B: agent is speaking C->>E: caller starts talking E->>E: voice-activity detector fires E->>C: clear buffered agent audio (0 RTT) E-->>B: control · speech_start B->>B: pacer.Barge(): drop carry, bump epoch B->>B: session: Speaking → Listening
The telephony edge holds the last few seconds of outbound audio and clears that buffer locally, so the caller stops hearing the agent before the interrupt reaches the bridge. The bridge then bumps a monotonic epoch on the pacer and drops the queued audio, so any text-to-speech frames still arriving from the abandoned turn carry a stale epoch and are discarded. It also moves the session from speaking back to listening and counts the barge-in.
The egress pacer
A synthesizer returns audio in chunks whenever it is ready. A phone line needs one 20 millisecond frame every 20 milliseconds. The pacer is the buffer between the two.
flowchart LR TTS["TTS audio
(arrives in bursts)"] --> Q["Pacer queue"] CLK["20ms tick"] --> DUE["Due(now)
drift-corrected"] Q --> DUE DUE -->|"one frame / period"| OUT["Caller"] DUE -->|"queue empty"| END["FlushTail → Listening"]
The detail that matters is drift. If the bridge sends a frame and waits 20 ms, scheduler jitter accumulates and the stream falls behind. Instead the pacer advances its target time by exactly one period per frame and resyncs forward if it stalls. Timing is tested by driving the pacer's clock with fixed values rather than sleeping, which is possible because the clock is injected.
End of playback
Generation finished is not playback finished. The model can produce the last token a second before the caller hears the last syllable. The media protocol has no playback-complete event, but because the pacer emits audio in real time, one frame per 20 ms, the bytes it has sent are a clock: N frames sent means N times 20 milliseconds of audio has left for the caller. Byte-accounting the paced output gives an accurate end-of-playback signal with no acknowledgement channel.
Endpointing and code-switching
A speaker mixing Bengali and English pauses while switching, mid-sentence. A silence timer short enough to feel responsive will cut that speaker off, so endpointing uses a semantic turn detector: the pause triggers an end-of-turn prediction, and only a confident result hands control to the agent. When the detector is unavailable the bridge falls back to its own silence timer rather than dropping the call. This tuning is validated against real recordings, because it is the main quality lever for a bilingual voice agent.
Part 03
The bilingual cascade
A managed speech-to-speech model takes audio in and returns audio out in one step, with lower latency and fewer parts. For English callers it is a reasonable default. Our callers speak Bengali and switch into English mid-sentence, and the managed options do not support Bengali on a telephony path. So the pipeline runs three stages: transcribe, generate, synthesize.
flowchart LR MIC["Caller
mu-law 8k"] --> STT["STT
realtime · bn + en"] STT -->|"transcript"| LLM["LLM
(the turn runner)"] LLM -->|"reply tokens"| RT{"RouteTTS(lang)"} RT -->|"bn"| AZ["Bengali neural TTS"] RT -->|"en"| CA["English TTS"] AZ --> OUT["mu-law 8k
→ caller"] CA --> OUT
The cost is one or two extra network hops. Two things keep them cheap: reply tokens stream from the model into the synthesizer, so audio starts before the sentence is complete, and the synthesizers emit mu-law directly, so their output goes into the pacer with no conversion.
Per-language routing
The best Bengali voice and the best English voice come from different vendors. The synthesis stage is a router: a small function maps a language to a driver (Bengali to the Bengali voice, everything else to the English voice), and a routing synthesizer wraps both behind one interface. The turn loop asks to speak text in a language and does not know which vendor answered. Each leg is built only if its credentials are present, so a single-language deployment runs without the other vendor.
One interface per vendor
flowchart TB CALL["Turn loop"] --> IF["Synthesizer interface
Open · PushText · Audio · Close"] IF --> FAC["Factory (name → driver)"] FAC --> A["azure"] FAC --> C["cartesia"] FAC --> R["routing (bn→azure / else→cartesia)"] FAC --> L["logged (keyless stub)"]
Both interfaces are token-in and bytes-out, which keeps first-audio latency low. A logged stub implements them with no credentials, so the whole pipeline runs on a laptop with silent audio. A fallback wrapper composes a primary transcriber with a backup: if the primary cannot start a session, the call uses the backup for its duration. It is configured as a piped choice, so operations can set a primary and a fallback without a code change.
Part 04
Reusing the chat engine
The voice bridge is a separate process at the edge. The chat intelligence (prompt assembly, retrieval, journeys, guardrails) lived inside the web server, in a large dispatcher a standalone binary cannot import. Giving voice its own simpler brain would drift from the chat brain over time, so there is one brain, reached over a small interface both sides can depend on.
flowchart TB
subgraph EDGE["Edge binary (cmd/voicebridge)"]
VB["Turn loop"]
end
subgraph SEAM["turnpipeline · stdlib-only, importable"]
RUN["Runner interface"]
SINK["Sink (streaming callbacks)"]
end
subgraph SERVER["Web server (has the chat engine)"]
ADP["Runner adapter"]
DISP["Chat dispatcher"]
end
VB -->|"Run(req, sink)"| RUN
ADP -. implements .-> RUN
ADP --> DISP
DISP -->|"deltas · status · result"| SINK
The interface is a request (text, language, channel), a streaming sink (callbacks for reply tokens, status, and the final result), and a runner with one method. Because the package is standard-library only, the edge imports it without pulling in the database, the model router, or the web framework.
Additive fan-out
The dispatcher already streams to the web widget through instance-level hooks keyed by session id. The reply text is already accumulated for every channel, so returning a result needs no new plumbing. The dispatcher gains a per-session registry of sinks, and each place it already emits a token or a status frame now also forwards to a registered sink, if one exists, after doing what it did before.
flowchart LR EMIT["dispatcher emit
(delta / status)"] --> HOOK["widget hook
(unchanged path)"] EMIT --> REG{"sinkFor(session)"} REG -->|"nil · all chat traffic"| NOOP["no-op"] REG -->|"a voice turn"| SINK["registered sink → TTS"]
The dispatcher is the busiest path in the platform, so a regression there is a platform incident. The safety property is that the sink lookup returns nil for every turn that is not a live voice call, so control flow and output for all existing traffic are identical. The change ships dormant until the voice adapter registers a sink, which is why it can land as a normal build, test, and commit ahead of any deployment. Wiring the real runner into a co-located edge is a separate, deliberate step.
Related reading
The same driver-and-factory pattern shows up across the platform. See the Common Channel Wrapper for the messaging-channel version of the same idea.
// RFC 2232 · Voice AI
// edit, open a PR, ships on merge to main