Skip to main content
Omazy Engineering

RFC 2232 · Voice AI

Voice AI

How a phone call runs as a real-time turn loop. This article covers the four parts of the system: the media bridge, turn-taking and barge-in, the bilingual speech cascade, and reusing the chat engine behind an importable interface.

// media bridge · barge-in · cascade · one brain // RFC 2232

Overview

The pipeline

A voice call has a hard timing constraint. The gap between the caller finishing a sentence and the agent starting to reply has to stay short, roughly under a second and a half at the median, or the caller talks over the agent. The system keeps that number low while reusing the same conversation engine as chat.

flowchart LR
  subgraph EDGE["Edge · co-located"]
    VOX["Telephony POP"]
    BR["Media bridge"]
    VOX <-->|"mu-law / 8k"| BR
  end
  subgraph CASCADE["Cascade (off-box)"]
    STT["STT"]
    LLM["LLM · chat engine"]
    TTS["TTS"]
  end
  subgraph WARM["Platform · warm path"]
    DB["chat_sessions
call_records"] end BR --> STT --> LLM --> TTS --> BR BR -.->|"signed control webhook"| DB
The whole system. A small co-located bridge moves mu-law audio between the telephony network and the speech and language models, and reports call lifecycle to the platform on a warm path.

Part 01

The media bridge

The bridge is a small, stateless Go service (cmd/voicebridge) that runs close to the telephony point of presence. It terminates the media WebSocket and runs the codec, the jitter buffer, the pacer, and the barge-in state machine. The expensive work runs elsewhere.

flowchart LR
  subgraph EDGE["Edge · co-located"]
    VOX["Telephony POP
SIP / PSTN + call scenario"] BR["Media bridge
cmd/voicebridge"] VOX <-->|"mu-law / 8k over WSS"| BR end subgraph BRAINS["Off-box vendors"] STT["STT"] LLM["LLM"] TTS["TTS"] end BR --> STT --> LLM --> TTS --> BR
Fig 1. The audio loop is co-located and small. Speech-to-text, the language model, and text-to-speech are network services the bridge streams to and from.

Co-location is a requirement, not an optimisation. The language model is the last stage of the pipeline, so any round trip added there is added to every turn. A bridge running far from the telephony POP spends its budget on network latency before the model produces a token. The bridge runs next to the POP; the durable platform can run anywhere because it is never on the audio path.

Mu-law end to end

Telephony audio is G.711 mu-law at 8 kHz. The bridge keeps that format for the whole round trip: caller audio arrives as mu-law, the transcriber accepts mu-law, the synthesizer is asked to return mu-law, and it goes back to the caller as mu-law. Skipping transcoding removes CPU work and latency. The one remaining per-packet cost is the base64 JSON framing the transport uses for each 20 millisecond packet, and the bridge uses pooled buffers so that framing adds no garbage-collection pressure. The frame protocol, the codec, a jitter-reorder buffer with packet-loss concealment, and the pacer live in one dependency-free package that is unit-tested without a socket.

The turn loop

A call is a sequence of turns. The bridge streams caller audio to speech-to-text. When speech-to-text reports the caller has finished, the bridge runs one turn and streams the reply into text-to-speech, whose audio the pacer sends back one frame at a time.

sequenceDiagram
  participant C as Caller
  participant B as Media bridge
  participant S as STT
  participant R as Turn runner
  participant T as TTS
  C->>B: audio frames (mu-law, 20ms)
  B->>S: stream audio
  S-->>B: final transcript
  B->>R: run one turn (text in)
  R-->>T: reply tokens (streamed)
  T-->>B: audio (mu-law)
  B->>C: paced 20ms frames out
Fig 2. One turn. Reply tokens stream into the synthesizer so audio can start before the full reply exists.

A session state machine tracks the call through listening, thinking, and speaking, and records the per-turn latency breakdown. The pacer owns the transition back to listening, which is covered in Part 02.

Part 02

Turn-taking and barge-in

Turn-taking needs two separate signals, and conflating them is a common mistake. One says the caller has started speaking, so the agent should stop (barge-in). The other says the caller has finished, so it is the agent's turn (endpointing).

Barge-in

A voice-activity detector. When there is speech energy while the agent is talking, stop the agent. It has to be instant, so it runs at the edge, next to the audio.

Endpointing

Deciding the caller is done is a judgement call. A naive silence timer cuts people off mid-thought. A semantic turn detector asks whether the utterance sounds finished, which matters for code-switching speakers who pause while changing language.

Barge-in at the edge

sequenceDiagram
  participant C as Caller
  participant E as Edge (VAD)
  participant B as Media bridge
  Note over C,B: agent is speaking
  C->>E: caller starts talking
  E->>E: voice-activity detector fires
  E->>C: clear buffered agent audio (0 RTT)
  E-->>B: control · speech_start
  B->>B: pacer.Barge(): drop carry, bump epoch
  B->>B: session: Speaking → Listening
Fig 3. The detector fires at the edge and drops the buffered agent audio immediately. The bridge is told after the fact and cleans up its own state.

The telephony edge holds the last few seconds of outbound audio and clears that buffer locally, so the caller stops hearing the agent before the interrupt reaches the bridge. The bridge then bumps a monotonic epoch on the pacer and drops the queued audio, so any text-to-speech frames still arriving from the abandoned turn carry a stale epoch and are discarded. It also moves the session from speaking back to listening and counts the barge-in.

The egress pacer

A synthesizer returns audio in chunks whenever it is ready. A phone line needs one 20 millisecond frame every 20 milliseconds. The pacer is the buffer between the two.

flowchart LR
  TTS["TTS audio
(arrives in bursts)"] --> Q["Pacer queue"] CLK["20ms tick"] --> DUE["Due(now)
drift-corrected"] Q --> DUE DUE -->|"one frame / period"| OUT["Caller"] DUE -->|"queue empty"| END["FlushTail → Listening"]
Fig 4. The pacer meters bursty TTS output into a steady 20ms cadence, and signals end of utterance when its queue drains.

The detail that matters is drift. If the bridge sends a frame and waits 20 ms, scheduler jitter accumulates and the stream falls behind. Instead the pacer advances its target time by exactly one period per frame and resyncs forward if it stalls. Timing is tested by driving the pacer's clock with fixed values rather than sleeping, which is possible because the clock is injected.

End of playback

Generation finished is not playback finished. The model can produce the last token a second before the caller hears the last syllable. The media protocol has no playback-complete event, but because the pacer emits audio in real time, one frame per 20 ms, the bytes it has sent are a clock: N frames sent means N times 20 milliseconds of audio has left for the caller. Byte-accounting the paced output gives an accurate end-of-playback signal with no acknowledgement channel.

Endpointing and code-switching

A speaker mixing Bengali and English pauses while switching, mid-sentence. A silence timer short enough to feel responsive will cut that speaker off, so endpointing uses a semantic turn detector: the pause triggers an end-of-turn prediction, and only a confident result hands control to the agent. When the detector is unavailable the bridge falls back to its own silence timer rather than dropping the call. This tuning is validated against real recordings, because it is the main quality lever for a bilingual voice agent.

Part 03

The bilingual cascade

A managed speech-to-speech model takes audio in and returns audio out in one step, with lower latency and fewer parts. For English callers it is a reasonable default. Our callers speak Bengali and switch into English mid-sentence, and the managed options do not support Bengali on a telephony path. So the pipeline runs three stages: transcribe, generate, synthesize.

flowchart LR
  MIC["Caller
mu-law 8k"] --> STT["STT
realtime · bn + en"] STT -->|"transcript"| LLM["LLM
(the turn runner)"] LLM -->|"reply tokens"| RT{"RouteTTS(lang)"} RT -->|"bn"| AZ["Bengali neural TTS"] RT -->|"en"| CA["English TTS"] AZ --> OUT["mu-law 8k
→ caller"] CA --> OUT
Fig 5. Transcribe, generate, synthesize. Audio is mu-law at 8 kHz at both ends, so nothing resamples.

The cost is one or two extra network hops. Two things keep them cheap: reply tokens stream from the model into the synthesizer, so audio starts before the sentence is complete, and the synthesizers emit mu-law directly, so their output goes into the pacer with no conversion.

Per-language routing

The best Bengali voice and the best English voice come from different vendors. The synthesis stage is a router: a small function maps a language to a driver (Bengali to the Bengali voice, everything else to the English voice), and a routing synthesizer wraps both behind one interface. The turn loop asks to speak text in a language and does not know which vendor answered. Each leg is built only if its credentials are present, so a single-language deployment runs without the other vendor.

One interface per vendor

flowchart TB
  CALL["Turn loop"] --> IF["Synthesizer interface
Open · PushText · Audio · Close"] IF --> FAC["Factory (name → driver)"] FAC --> A["azure"] FAC --> C["cartesia"] FAC --> R["routing (bn→azure / else→cartesia)"] FAC --> L["logged (keyless stub)"]
Fig 6. Transcription and synthesis each sit behind a narrow streaming interface resolved from a name-keyed factory. Adding or replacing a vendor is one implementation file and one factory case.

Both interfaces are token-in and bytes-out, which keeps first-audio latency low. A logged stub implements them with no credentials, so the whole pipeline runs on a laptop with silent audio. A fallback wrapper composes a primary transcriber with a backup: if the primary cannot start a session, the call uses the backup for its duration. It is configured as a piped choice, so operations can set a primary and a fallback without a code change.

Part 04

Reusing the chat engine

The voice bridge is a separate process at the edge. The chat intelligence (prompt assembly, retrieval, journeys, guardrails) lived inside the web server, in a large dispatcher a standalone binary cannot import. Giving voice its own simpler brain would drift from the chat brain over time, so there is one brain, reached over a small interface both sides can depend on.

flowchart TB
  subgraph EDGE["Edge binary (cmd/voicebridge)"]
    VB["Turn loop"]
  end
  subgraph SEAM["turnpipeline · stdlib-only, importable"]
    RUN["Runner interface"]
    SINK["Sink (streaming callbacks)"]
  end
  subgraph SERVER["Web server (has the chat engine)"]
    ADP["Runner adapter"]
    DISP["Chat dispatcher"]
  end
  VB -->|"Run(req, sink)"| RUN
  ADP -. implements .-> RUN
  ADP --> DISP
  DISP -->|"deltas · status · result"| SINK
Fig 7. A stdlib-only package holds the interface. The edge imports it; the server implements it by delegating to the existing dispatcher. Neither side imports the other.

The interface is a request (text, language, channel), a streaming sink (callbacks for reply tokens, status, and the final result), and a runner with one method. Because the package is standard-library only, the edge imports it without pulling in the database, the model router, or the web framework.

Additive fan-out

The dispatcher already streams to the web widget through instance-level hooks keyed by session id. The reply text is already accumulated for every channel, so returning a result needs no new plumbing. The dispatcher gains a per-session registry of sinks, and each place it already emits a token or a status frame now also forwards to a registered sink, if one exists, after doing what it did before.

flowchart LR
  EMIT["dispatcher emit
(delta / status)"] --> HOOK["widget hook
(unchanged path)"] EMIT --> REG{"sinkFor(session)"} REG -->|"nil · all chat traffic"| NOOP["no-op"] REG -->|"a voice turn"| SINK["registered sink → TTS"]
Fig 8. The widget path is untouched. The new branch is guarded by a registration that only a voice turn performs, so for every chat, widget, and API turn the lookup returns nil and the branch does nothing.

The dispatcher is the busiest path in the platform, so a regression there is a platform incident. The safety property is that the sink lookup returns nil for every turn that is not a live voice call, so control flow and output for all existing traffic are identical. The change ships dormant until the voice adapter registers a sink, which is why it can land as a normal build, test, and commit ahead of any deployment. Wiring the real runner into a co-located edge is a separate, deliberate step.

Related reading

The same driver-and-factory pattern shows up across the platform. See the Common Channel Wrapper for the messaging-channel version of the same idea.

// RFC 2232 · Voice AI

// edit, open a PR, ships on merge to main