Skip to main content
Omazy Engineering

RFC 2222 Part V · voice configuration

Authored top down. Resolved bottom up.

A voice agent is configured in one direction and answers the phone in the other. Reading that configuration once at process start freezes the authoring order at boot, and every defect that followed descended from the inversion: one process serving one tenant, a console writing a table nothing read, and cached audio that outlived the words it was made from.

// per-call resolution · content-hashed cache · routing before the model // the sequel to the latency post, and a different subject
3.50 s
a generated turn: model plus synthesis
0.75 s
the same answer routed, cold cache
0.16 ms
routed with the phrase already warm
12.2 ms
of synthesis per character, measured
3 / week
voice changes per language before an override

Section 01 · the spine

Configuration is authored in one direction and consumed in the other.

Someone setting up a phone agent works downward. They pick the agent, then the languages it speaks, then the voice each language is spoken in, then the fixed phrases and the answers and the tuning. That is the order the console is laid out in and the order a person thinks in.

A call runs the same chain in reverse. A number rings, and only then is there an app. From the app comes the profile, from the profile the language, from the language the voice, from the voice the greeting, and only after all of that does the first turn happen. Nothing in that chain is knowable before the call arrives, because the first link is the dialed number.

flowchart TB
  subgraph AUTH["AUTHORING · top down, in a console"]
    direction TB
    A1["agent"] --> A2["voice profile"] --> A3["language"] --> A4["voice"] --> A5["phrases, answers, tuning"]
  end
  subgraph RUN["RUNTIME · bottom up, per call"]
    direction TB
    R1["inbound call"] --> R2["app, from the dialed number"] --> R3["profile"] --> R4["language"] --> R5["voice"] --> R6["greeting"] --> R7["per turn: route or generate"]
  end
  BOOT["process start reads the environment"] -.->|"freezes the authoring order at boot"| RUN
Fig 1. The two orders. Reading configuration once at process start does not resolve the runtime chain, it pins the authoring chain to whatever the process was started with, which is a decision made before any caller existed.

The original implementation read all of it from environment variables at process start. That is not a shortcut on the runtime path, it is the authoring path frozen at boot, and the consequences fall out mechanically. One process served exactly one tenant. In one language. With one voice. A table the console wrote to was read by nothing at call time, so an operator could save a greeting, see it stored, and hear a different greeting when they dialled.

Worth being precise about the shape of the bug: nothing was broken in the sense of throwing. Every layer did what it was told. The inversion just made the thing being told to it belong to a different tenant, and there is no error code for that.

Section 02 · the checklist

A half-built version ships if any one of five is left open.

Every one of these is individually plausible, individually defensible in review, and individually enough to make the whole feature a lie. They were written down before the work started, as a list of ways to appear finished.

#The leakWhat closes it
1 Configuration lives in the database, the process reads environment variables at boot. Resolve per call, from the app the call is for. Environment stays as the fallback, so an app that was never migrated behaves byte for byte as it did.
2 The cache is keyed by voice, but the list of phrases to warm is built from the environment. Drive the warm from the resolved configuration, and re-run it when the resolved fingerprint changes.
3 The greeting lives in one system and the acknowledgements live in another. Make every fixed spoken phrase one object with a kind. One table, one cache, one warm, one editor, one usage counter.
4 A rate limit is enforced in the user interface. Enforce it in the service on the write path. The console and the tool API both reflect it rather than each implementing it.
5 A console reports a line as live because a row exists. Compute readiness from what the call path actually consumes, and report an unknown as unknown.

The fifth had already happened once on this surface before the list existed, and cost a day. A console reported a line as live, an operator believed it, and the call path was reading something else entirely. That is the reason readiness is computed rather than displayed, which Section 13 gets to.

All five share one property: each can be closed with a change that makes the system look identical from the outside and behave completely differently. That is what makes them worth writing down in advance, because none of them announces itself in a demo.

Section 03 · leak one

Resolve when the call arrives, and keep the old path as the fallback.

The fix is one function called per call rather than per process. It answers a single question: what should this app sound like on this call. Languages in spoken order, a voice per language, what the recognizer listens for, the fixed phrases, the tuning knobs, and the cache namespace they all hash to.

flowchart TB
  CALL["call arrives, dialed number in the start frame"] --> APP["resolve the app"]
  APP --> GET["Resolve: profile rows plus the phrase set"]
  GET --> D1{"profile row found?"}
  D1 -->|"no"| ENV["environment configuration
byte for byte unchanged"] D1 -->|"yes"| D2{"enabled?"} D2 -->|"no · staged for a cutover"| ENV D2 -->|"yes"| OVER["overlay languages, voices, hints,
tuning, phrases, routed answers"] ERR["lookup error, or the 2s timeout"] --> ENV OVER --> NS["namespace = hash over profile AND phrase text"] NS --> WARM["background warm, once per namespace per process"] NS --> SPEAK["greeting, then per turn: route or generate"]
Fig 2. Per-call resolution with the environment as the fallback. Four different outcomes converge on the same branch: no row, a staged row, a lookup error, and a lookup that ran out of time. Only an enabled profile changes anything about the call.

The fallback is the part that made this shippable. A live line was already answering calls from environment configuration, so an app with no profile row has to keep behaving byte for byte: same greeting, same acknowledgements, same recognizer hints, same tuning. Treating the fallback as a transitional courtesy would have made the migration a cutover for everyone at once. Treating it as a permanent, tested branch made it a per-app decision.

// Every failure mode returns the environment configuration unchanged, with a
// log line naming which one happened. "It fell back" must never be
// indistinguishable from "it was configured that way".
p, found, err := d.profiles.Resolve(lookupCtx, key)
switch {
case err != nil:
    log.Printf("voice profile lookup failed (%v) - using env configuration", err)
    return d
case !found:
    return d
case !p.Enabled:
    // A staged profile. The row exists so an owner can prepare a cutover;
    // until it is enabled the call path must not read it.
    log.Printf("voice profile is disabled - using env configuration")
    return d
}
return d.withProfile(p)

Staged is a real state

A profile row exists in a disabled state so an owner can author the whole thing, read it back, and cut over deliberately. Seeding a profile from the older configuration and enabling it in the same step would have silently changed the voice and the greeting on a line answering calls that day. The two sources already disagreed, which is exactly why nobody should find out from a caller.

The lookup is bounded

Two seconds, and then the environment answers. The greeting cannot start until the resolve returns, and a caller hearing nothing is a worse failure than a caller hearing a slightly generic greeting. A timeout that degrades is a design decision; a timeout that blocks the greeting is an outage with extra steps.

One detail that reads as a nicety and is not: an enabled profile with no greeting phrase does not fall silent. The overlay only replaces what the profile actually carries, because an empty phrase set overwhelmingly means "not authored yet" rather than "say nothing". Tuning is the deliberate exception and wins outright, zeros included, since a zero jitter lead and zero echo suppression are both real settings and treating zero as unset would make them unreachable.

Section 04 · leak three

A greeting and an acknowledgement are the same object.

A greeting, a filler said while the model thinks, a farewell, an idle nudge and a verbatim answer to a common question are all one thing: a fixed phrase, authored per language, delivered word for word. They differ only in what triggers them.

They started life in three different places, so each had its own editor, its own idea of what a language is, and its own failure mode. Unifying them is what makes everything downstream happen once instead of four times: one cache, one warm, one duration readout, one usage counter, and one place an owner edits the words their business says out loud.

-- Every fixed spoken phrase is the same object. They differ only in what
-- triggers them, so the trigger is the column, not the table.
ALTER TABLE message_templates
  ADD COLUMN kind     TEXT  NOT NULL DEFAULT '',
  ADD COLUMN triggers JSONB NOT NULL DEFAULT '[]'::jsonb;

ALTER TABLE message_templates ADD CONSTRAINT kind_check
  CHECK (kind IN ('', 'greeting', 'filler', 'farewell', 'nudge', 'answer'));

-- The spoken form is a SEPARATE column from the on-screen blocks, not a
-- rendering of them. What reads well on a screen is wrong out loud, and a
-- synthesizer reads the asterisks.
ALTER TABLE message_template_variants
  ADD COLUMN spoken_text TEXT NOT NULL DEFAULT '';

The trigger becomes a column, not a table. That single decision is what lets the routed answers ride the same per-call resolve as the greeting: by the time the profile is in hand, the answers and their trigger phrases are already there, so routing costs no query and no round trip on the turn path.

The spoken text is a separate column from the on-screen blocks, not a rendering of them. Copy that reads well on a screen is wrong out loud, and a synthesizer will happily read the asterisks. Same phrase, two renderings, authored independently. Deriving one from the other looks like deduplication and produces an agent that says punctuation.

Section 05 · the hard part

The cache key is a hash over the material, not a version counter.

Cached audio is correct exactly as long as the thing it was made from has not changed. The naive design gives the profile a version counter, bumps it on any voice change, and prefixes the cache key with it. That works for every change that goes through the voice editor, and only those.

The hole is the content API. A phrase edited through it changes what the agent says while knowing nothing about voice at all, so nothing bumps the counter, the key stays the same, and the old audio keeps playing. The agent is now saying the previous version of a sentence that no longer exists in the database, and there is no error anywhere.

// The cache namespace is a hash over everything that changes what a caller
// hears, not a bare version counter.
//
// A counter is not enough on its own. A phrase edited through the content API
// changes the words the agent says while knowing nothing about voice, so
// nothing bumps the counter and the old audio keeps playing. Hashing the
// material makes stale audio unreachable by construction rather than by
// someone remembering to invalidate it.
func (p Profile) fingerprint() string {
    h := sha256.New()
    write(h, p.AppID.String(), strconv.Itoa(p.Version), p.PrimaryLanguage)
    write(h, p.Languages...)
    for _, l := range p.Langs {
        write(h, l.Lang, l.VoiceID, l.FallbackVoiceID)
        write(h, l.STTHints...)
    }
    for _, ph := range p.Phrases {
        write(h, ph.Kind, ph.Lang, ph.Key, ph.Text) // the words are in the key
    }
    return hex.EncodeToString(h.Sum(nil))[:16]
}

Hashing the material closes it by construction. The namespace covers the app, the version, the language list, every voice id, the recognizer hints, and the text of every phrase. Any edit to any of those produces a different namespace, so previously cached audio is not invalidated, it is simply unreachable. Nobody has to remember anything.

Unreachable beats invalidated

An invalidation is a step someone can forget, a code path that can throw, and a race two processes can lose. A key derived from content is none of those: the reader computes the same namespace the writer did, from the same inputs, or it computes a different one and misses. The failure mode of a miss is a slow call, which is survivable.

Voice and language are in the key too

Below the namespace, the entry key still carries the voice, the language, the audio format and the sample rate alongside the text. The same words in a different voice are different audio, and serving one caller the voice a different caller picked is precisely the bug this must not have.

The same discipline is applied one layer up, to measurements rather than audio. A recorded duration keyed only to a phrase row outlives an edit of that phrase's words, so it carries a hash of the exact text it measured and is dropped when the text changes. Section 12 has the reasoning.

Section 06 · the same bug twice

A warm that fills a namespace nobody reads.

Both of these shipped. Both are the same shape: audio is rendered under one namespace and looked up under another. The cache fills, the log says warm, the metrics look healthy, and every single caller pays full synthesis.

One: the boot warm, after per-call resolution landed

The process warms its fixed phrases at startup. That was written before profiles existed, so it renders the environment phrase set under an empty namespace. For an app on the environment fallback that is exactly right, because its calls also run with an empty namespace. For an app with an enabled profile it is silently useless: those calls read under the profile fingerprint, so every boot-warmed entry sits under a key nothing will ever look up.

It cannot be fixed at boot either. The media process has no database, does not know which apps will ring it, and with per-call resolution the phrase set is not knowable until a call arrives. So the warm follows the profile instead: the first call that resolves a given namespace triggers one background pass, and every later call for that namespace is served from cache. Keyed by namespace rather than by app, so a voice or phrase edit re-warms and an unchanged profile never warms twice.

The cost is one cold call per profile version per process, which is exactly what the caller paid before any cache existed. Warming synchronously instead would put the entire phrase set in front of the first caller's greeting.

// This bug is invisible to a warm that merely succeeds, so the test asserts
// the NAMESPACE each phrase was rendered under, not that rendering happened.
// Against the old behaviour it fails with: warmed under namespace "".
for _, got := range fake.rendered {
    if got.Namespace != profile.CacheNamespace {
        t.Fatalf("phrase %q warmed under namespace %q, want %q",
            got.Text, got.Namespace, profile.CacheNamespace)
    }
}

This is the part worth keeping. A warm that merely succeeds proves nothing here, because succeeding is not the failure mode. The test asserts the namespace each phrase was rendered under, and it was verified against the old behaviour to make sure it actually fails: warmed under namespace empty string. A regression test that has never been seen red is a hypothesis.

Two: the change path, computing a namespace without phrases

The voice change service also warms. It read the profile, computed the namespace it would have after the change, and rendered every phrase under it. But it read the profile without its phrases, and the namespace is a hash over the profile and the phrase text. Different inputs, different hash, same class of dead warm.

// The change path resolves the profile exactly the way the CALL PATH sees it:
// rows PLUS the whole phrase set, with the namespace recomputed over both.
//
// The phrases are not an optimisation here. The namespace is a hash over the
// profile AND its phrase text, so a profile read without phrases produces a
// different namespace from the one the call path looks audio up by. Warming
// under that namespace succeeds, reports every phrase rendered, and leaves the
// call path cold: a failure that looks exactly like success.
func (s *Service) resolve(ctx context.Context, appID uuid.UUID, lang string) (Profile, []Phrase, string, error) {
    prof, err := s.profiles.Get(ctx, appID)
    if err != nil {
        return Profile{}, nil, "", err
    }
    all, err := s.phrases.SpokenPhrases(ctx, appID)
    if err != nil {
        return Profile{}, nil, "", err
    }
    // One shared mapping with the resolver. Two copies would produce two
    // namespaces, and the second one warms a key nothing ever reads.
    return prof.WithPhrases(PhrasesFrom(all)), phrasesForLang(all, lang), lang, nil
}

The fix is not a second call to add the phrases. It is one shared code path with the resolver, including one shared mapping function from stored rows to phrase values, because two copies of that mapping would eventually produce two namespaces. The test derives its expectation from the resolver rather than restating a namespace literal, so the test and the call path cannot drift apart either.

Generalised: when a value is a hash over inputs, the set of inputs is an interface, and every producer of that value has to be the same code as the consumer. Anything else is a duplicated definition of correctness that compiles.

Section 07 · what to cache, and what to drop

Caching everything is worse than caching nothing.

The obvious version of a synthesis cache stores all output. It is a trap. Generated replies are never byte-identical twice, so an everything-cache pays full storage for a hit rate near zero, and then evicts the handful of fixed phrases that would actually have hit.

So caching is opt-in per request. The caller marks a synthesis cacheable, and nothing else is stored. That is a two-line change with an outsized effect on the hit rate, and it also makes the cache's contents a set someone can reason about: it is the authored phrases, in the configured voices, and nothing else.

The eviction policy that was correct until it was not

The first version had no eviction at all, and the reasoning was sound at the time: the cacheable set is a small fixed list known at configuration time, so a least-recently-used policy could only ever discard something about to be needed again. There was a hard ceiling as a guard, and past it synthesis still worked, it just stopped storing.

A switchable voice invalidates the premise. The set is no longer fixed, it turns over, and roughly four voice changes fill the ceiling with renderings in voices nobody uses any more. The behaviour past the ceiling made it worse than a simple slowdown: one log line, then silence, while every phrase authored or re-voiced afterwards missed forever. The only evidence was a single line that had scrolled past hours earlier.

Least recently used, plus a tiebreak

Entries carry a monotonic counter alongside the wall-clock stamp. Warming runs in a tight loop and the clock can return the same instant for several entries, which would make eviction order arbitrary at exactly the moment the cache is under pressure. A counter makes it deterministic.

A retired voice is invalidated explicitly

Changing the namespace already makes the old audio unreachable. Invalidating the retired voice is a separate step for a separate reason: unreachable bytes still occupy the ceiling, and the phrases callers actually hear would be evicted to make room around them.

A capacity policy is only ever correct with respect to an assumption about the workload. Writing the assumption down next to the policy is what let this one be caught by reading rather than by an incident, because the feature that broke it was named in the same review.

Section 08 · determinism

Some answers should never reach the model.

Opening hours. The address. How to book. The fee policy. Every caller asks a handful of the same questions, the answers do not change, and a person has already written them down. Sending those through a language model buys nothing and costs the caller the full turn.

One turn, three paths · benchmark, 2026-08-11

generated 3.50 s
routed cold 0.75 s
routed warm 0.16 ms

generated 1.45 s model plus 2.05 s synthesis of a 148 character reply

routed cold no model, full synthesis of the 42 character phrase

routed warm no model, no synthesizer, a map lookup and a copy

Measured as a benchmark rather than a test, because paying three and a half seconds of real sleep on every run would be a poor trade for a number that does not change. A cold routed turn is already 4.6 times better than a generated one, and it pays that price once per phrase per voice.

On live calls the same shape holds with the network in it: the first call under a new namespace pays roughly three and a half seconds to first audio, and once the namespace is warm the same phrase lands around 850 milliseconds. The floor underneath that is transport, not software, which is the subject of the latency piece rather than this one.

The guardrail is the whole design

A confidently wrong verbatim answer is worse than a slow correct one. A slow answer costs a caller two seconds; a wrong one said with total confidence costs them the reason they called. So every rule in the matcher is biased towards handing the turn back to the model.

flowchart TB
  U["final transcript"] --> T{"any trigger matched?"}
  T -->|"no"| M["the model answers the turn"]
  T -->|"yes"| L{"one unambiguous winner?"}
  L -->|"no · two responses fired"| M
  L -->|"yes"| W{"utterance under the word cap?"}
  W -->|"no"| M
  W -->|"yes"| LOC{"spoken text in this language?"}
  LOC -->|"no"| M
  LOC -->|"yes"| A["speak it verbatim, from cache"]
  M -.->|"every refusal is logged with the candidate and the reason"| LOG["operator log"]
Fig 3. Four ways to refuse and one way to route. Every refusal path ends at the model, which is allowed to be slower and is allowed to answer two questions at once.
GuardrailSettingWhy
minimum trigger length3 code pointsA two letter trigger matches a large share of any language, and nobody writing a one word trigger means "fire on every call".
word boundary, left onlyrequiredSo a trigger for booking does not fire inside a longer word. The right edge is deliberately unchecked, because Bengali attaches case endings to the noun and demanding a right boundary would switch the feature off in the language it mostly answers in.
single unambiguous winnerlongest, containingThe longest matched trigger must contain every other match. A longer specialisation of the same phrase wins; two unrelated matches mean the caller asked two things, so the model answers.
utterance length cap16 wordsA verbatim answer replies to exactly one question. Anything discursive almost always carries a second clause, and answering half of it confidently is the failure this is arranged to avoid.
locale fallbackrefuseNo spoken text in the language of the call means the model answers. Replying in the wrong language is worse than replying slowly.

Refusals are logged with the candidate that nearly fired and the reason it did not, and only when something nearly fired. Logging every ordinary turn as "no trigger matched" would bury the interesting line. This matters because "the agent said the wrong thing" and "the agent ignored my saved answer" arrive as the same sentence from an operator, and only the reason separates them.

Routing is fed only from an enabled profile. An app still on the environment fallback has no answers at all, so the feature is switched off for it by construction rather than by a flag someone has to set correctly. The safest way to not leak a new behaviour into old tenants is to make the new behaviour depend on data they do not have.

Section 09 · the cost model

Two different curves, and the bug that comes from conflating them.

Synthesis time is linear in characters: a fixed cost of 0.24 seconds plus 12.2 milliseconds per character, measured on real replies. How long a caller hears the result is a completely different curve, roughly 15.4 characters per second. Both are used all over the product, and they are not interchangeable.

Both constants live in one place, in one package, for one reason: two features need the same number and must not each invent their own. The impact preview quotes it to an owner before they commit, and the latency work measures itself against it. A second copy would drift the moment either is retuned. The console imports the same two constants rather than restating them, and the one place a literal was hardcoded instead is the bug at the end of this section.

Same text, two answers · derived from the two constants

synthesis time audible length
a saved answer, written to be heard 42 chars

synthesis 0.75 s · audible 2.7 s

the mean authored phrase 110 chars

synthesis 1.58 s · audible 7.1 s

a generated reply of average length 148 chars

synthesis 2.05 s · audible 9.6 s

a list-shaped answer the model chose to give 317 chars

synthesis 4.11 s · audible 20.6 s

where 12 seconds of synthesis lands 964 chars

synthesis 12.00 s · audible 62.6 s

Both bars on each row are on the same scale. The gap between them widens with length, which is why a threshold set on one curve is meaningless on the other.

Count code points, not bytes

The vendor meters characters. A Bengali character is three bytes in UTF-8, so counting bytes triples every estimate for exactly the tenant this was measured on. This is a one-line detail that turns a cost model into a fiction, and it has to be got right in the Go and in the browser independently, because the editor prices a phrase as it is typed.

The conflation, with numbers

The console warns an author when a phrase is too long to say to a caller, at 12 seconds. That threshold belongs on the audible curve. Put it on the synthesis curve by mistake and it does not fire until 964 characters, which is 63 seconds of audio: nobody is listening by then, and the warning designed to prevent exactly that has stayed quiet the whole time.

A 12 second threshold, put on the wrong curve

intended 12.0 s
actual 63 s

Seconds of audio a caller would sit through before the warning appeared. The real defect that shipped was smaller and the same species: one surface highlighted past a hardcoded literal while the warning next to it read the shared constant, so retuning one would have silently left the other behind.

A number that appears twice in a codebase is a number that will eventually appear twice with different values. The cheapest fix is that the second occurrence imports the first, and the second cheapest is a test that fails when they disagree.

Section 10 · leak four

Changing a voice is a priced operation, not a dropdown.

Picking a different voice invalidates every cached phrase for that language and voice. The next callers pay full synthesis until the cache refills, the change is heard by everyone immediately, and a few of them in a row fill the cache ceiling with audio nobody will ever hear again. None of that is visible from a select element.

sequenceDiagram
  participant O as Owner
  participant S as Change service
  participant C as Phrase cache
  participant P as Callers on the line
  O->>S: preview, language plus candidate voice
  S-->>O: phrase count, seconds of audio, seconds of synthesis, allowance left
  O->>S: confirm
  Note over S,C: the profile still points at the OLD voice
  S->>C: render every phrase under the NEXT namespace
  C-->>S: per phrase: rendered, measured, or failed with a reason
  P->>C: calls in flight keep hitting the OLD namespace
  S->>S: write voice and version in one transaction
  S->>C: invalidate the retired voice
  P->>C: the next call resolves the NEXT namespace, already warm
Fig 4. The order is the design. The new voice is rendered under the namespace the profile will have after the cutover, while the profile still points at the old one, so calls in flight keep hearing the same person for the whole of the warm.

Quantified before, not reassured during

The preview is computed from the real phrase set: how many phrases, how many seconds of audio, how many seconds of synthesis, and the wait one caller actually feels, which is the average phrase rather than the total. "This may take a moment" is not a number an owner can decide with.

Progress per phrase, not one aggregate

Twenty-three of twenty-four succeeded tells an owner nothing about which line their callers will not hear. One phrase failing does not stop the rest either: a single unsayable line is a line to fix, and refusing to warm the others over it would leave every caller paying for all of them.

All of them failing is a refusal

If every phrase fails to render, the overwhelmingly likely cause is a voice id the provider will not accept, and cutting over to it would answer the phone with silence. The change is refused rather than applied. Silence is the one failure a caller cannot work around.

The ledger is written last

A failed change never spends the week's allowance. The write is also best effort: the voice has genuinely changed by then, and refusing to report that because the audit insert failed would leave someone retrying a change that already happened.

The limit lives in the service, and it is not a wall

Three changes per app per language per week, counted from the audit ledger and enforced on the write path, where every caller has to pass through it. Enforcing it in the console would make it a suggestion the moment a second client exists, which is Section 11.

A hard limit would be wrong, though, and the reason is worth stating: when the voice itself is the defect, when it mispronounces the business name or reads a language badly, it has to be fixable immediately. So an owner can override, and the override requires a written reason and is logged with it. An unexplained override is just a higher limit.

A rate limit with no escape hatch converts a quality problem into a week-long outage. A rate limit with an unlogged escape hatch converts itself into decoration. The reason field is what makes the third option a real one.

Section 11 · the second client

Configuring the whole thing from a tool, without a second set of rules.

Everything else about an account can already be set up over MCP. Voice was the one surface an operator or an assistant could not touch at all, which meant an account could be configured end to end except for the part that answers the phone.

The tools are a thin layer over the same service the console calls. That is the entire design constraint, and two things follow from it that are not stylistic.

The voice change tool is two-step

Without an explicit confirm it changes nothing and returns the priced impact plus the remaining allowance. With confirm it goes through the same service, the same warm order, the same limit. So a tool caller cannot bypass a warning the interface would have shown, and there is one rate-limit counter rather than two implementations of one.

Every write returns readiness

Computed from what the call path consumes, not from what was saved. This surface has already shipped one configured-but-not-callable trap. A tool that answers "saved" without answering "callable" is how the second one ships, and an assistant will believe "saved" indefinitely.

There is a smaller lesson underneath. Once a second client exists, every rule that was implemented in the first client is a rule that does not exist. The useful test is not whether the two clients agree today, it is whether disagreeing is possible: if the number can only come from one place, it cannot.

Section 12 · honesty about numbers

Estimated, then measured, and one of them never gets to be.

Every duration in this system starts as an estimate off the length of the text, and becomes a measurement only after a real synthesis has actually happened. Anything shown to a person carries a flag saying which it is, and the flag flips on its own the first time a phrase is genuinely rendered.

// A measurement keyed only to a phrase row outlives an edit of that phrase's
// words: rewrite the greeting and the row still reports the old length. So the
// row carries a hash of the exact text it measured, and the reader drops any
// row whose hash no longer matches the phrase as it stands today.
//
// "Measured" that is silently stale is a worse lie than "estimated" was,
// because it sounds authoritative.
type PhraseAudio struct {
    VariantID  uuid.UUID
    VoiceID    string
    Format     string
    SampleRate int
    Bytes      int
    DurationMS int
    TextSHA    string // exact: no trimming, no case folding, no normalising
}

The interesting failure is not an estimate being wrong, it is a measurement going stale. A duration keyed only to a phrase row outlives an edit of that phrase's words: rewrite a greeting, and the row still reports the old length for words nobody says any more. So it carries a hash of the exact text it measured, and the reader drops any row whose hash no longer matches. Exact, with no trimming and no case folding, because a leading space is a different synthesis request and may well be different audio.

A stale row is dropped rather than returned with a caveat. A caller that receives a duration will report it, and the caveat does not survive the trip to a console. That is the same argument as the cache namespace, applied one layer up: make the wrong answer unreachable rather than annotated.

What can become measured

Audio length. The rendered format is fixed-rate, so the byte count is the duration, with no guessing involved. A compressed format is not, and the function returns unknown rather than a plausible number, because guessing is worse than saying so.

What never can

Time saved. Synthesis avoided by routing an answer is a counterfactual: it is a spoken count multiplied by an estimator, and no amount of real data turns it into a measurement. It is labelled estimated permanently, because a number that sounds authoritative and is structurally unknowable is the worst kind to print.

The rule this settles on: "measured" that is silently stale is a worse lie than "estimated" ever was, because it sounds authoritative. Any label asserting provenance needs its own invalidation, or it is decoration with a strong opinion.

Section 13 · leak five

Configured is not the same as callable.

The readiness rubric is a list of checks, and every one is phrased against what the call path actually consumes rather than against whether somebody saved something. That framing is the whole feature.

CheckPhrased asWhy the obvious version lies
enabled The call path reads this profile A staged profile is not a mild warning. Every call is answered from the process configuration, which belongs to whichever tenant that process was started for.
greeting A greeting phrase per language A missing greeting is not silence. The overlay only replaces what the profile carries, so the older greeting keeps playing: another business's words, in this business's voice.
voices Reported unknown, not done An empty voice id is not broken, it means the media process falls back to its own default, and this process cannot see what that is. Reporting done would be the same lie in a smaller font.
fillers Off is a valid state, reported Explicitly off means callers hear nothing while the model thinks. That is a legitimate choice and a surprising one, so it is shown rather than hidden as a pass.
phone_number A number routes to this app Zero active numbers means nobody can dial it, however complete everything else is. When the check cannot run at all it reports unknown, never a pass.

Only one field means what an operator wants to know, and it is true only when nothing blocking is outstanding: a caller dialling this number right now is answered by this configuration. Everything else is detail underneath it. The summary sentence when it is false names the blockers and says the quiet part: configured is not the same as callable.

Three statuses were not enough. Done, missing and pending all imply the check ran; unknown is for the cases where this process genuinely cannot see the answer, and collapsing it into either of the others is how the original trap worked.

Implementation surface

Where it lives.

Four small packages and a schema change, plus the two clients that both call through them rather than around them.

Resolution and phrases

  • internal/voiceprofile: the per-call resolve, the content-hashed namespace, the readiness rubric
  • internal/messagetemplates: one phrase table with a kind, the spoken text column, the measurement rows
  • internal/voicebridge: the environment fallback, the overlay, the per-namespace background warm

Determinism and change

  • internal/voiceroute: anchored, single-winner matching with logged refusals
  • internal/voicetone: priced impact, ordered re-warm, invalidation, the enforced allowance
  • internal/speech: the opt-in cache with least-recently-used eviction, and the one copy of the cost model
  • internal/adminmcp: the two-step tools, thin over the same service

Related reading

The latency side of the same stack, the silent-call bugs, the measured budget, the phrase cache and streaming synthesis, is From a silent call to 820 milliseconds. The architecture underneath both, the media bridge, the pacer and the bilingual cascade, is Voice AI. The driver-and-factory shape every vendor sits behind is described for messaging in the Common Channel Wrapper, and the per-turn provenance operators read is Conversation Observability.

// RFC 2222 Part V · voice configuration and determinism

// edit, open a PR, ships on merge to main