Skip to main content
Omazy Engineering

field notes · model routing

One prompt, three models.

A billing failure took two model providers offline at once and moved a whole agent fleet onto its last fallback without a single failed deploy. Getting back out produced measurements for three model families against one strict system prompt, and almost none of them matched how those models looked on a short test prompt.

// measured on production traffic · one agent, one prompt, three families // every figure came off production traffic, not a benchmark
81%
of every reply was spent on providers that could not answer
22×
latency difference between a probe prompt and a production one
0 of 3
models needed no prompt change to behave the same way
470
of 555 output tokens spent thinking, on one short answer
41.3
words per sentence, satisfying a rule written in sentences

Section 01 · the short version

A model swap is a prompt change you did not write.

Two model providers stopped serving on the same day because their accounts ran out of money. Nothing failed to deploy, nothing paged, and no request errored from the outside. Traffic simply slid down a fallback chain onto a different model, and the agents kept answering with a different personality, different reliability, and different failure modes.

The interesting part was not the outage. It was that a fleet of agents carefully tuned against one model behaved measurably differently on the next one down the chain, in ways nobody had written a test for, and that the swap itself was invisible in every dashboard except latency. What follows is what we would want to have known first: the classification bug that made a permanent failure look retryable, the arithmetic of a fallback chain with no circuit breaker, the reason a passing preflight still let through a model four times slower, and the behavioural differences between three model families running one strict system prompt.

Section 02 · classification

Running out of money looks exactly like being rate limited.

Every retry policy we have ever written keys off the status code. That works until the status code is chosen by a billing system rather than a capacity one, and the two disagree about what the caller should do next.

Here is the same condition, an account with no remaining credit, as reported by three different layers of the same request path:

Reported byStatusWhat the body saidWhat a status-only policy does
Vendor A, directly 429 insufficient_quota retries the full ladder
Vendor B, directly 400 credit balance is too low fails fast, correctly, by accident
Gateway, in front of vendor A 503 upstream provider unavailable retries, and the body says nothing

A 429 that means "you are going too fast" and a 429 that means "your card declined" are the same integer. The first is worth waiting out. The second will still be true in an hour, and every retry against it is pure latency charged to a user who is watching a typing indicator. The fix is to classify on the response body, and to keep the marker list deliberately narrow.

// Status is the thing that misleads: one vendor bills credit exhaustion
// as 429, another as 400, and a gateway in front of either can report 503
// with no billing hint at all. Classify on the body, and keep the list
// narrow — a plain per-minute rate-limit message IS worth retrying.

var permanentMarkers = []string{
    "insufficient_quota",
    "credit_balance_exhausted",
    "no credits remaining",
    "credit balance is too low",
    "billing_hard_limit_reached",
    "exceeded your current quota",
    "invalid_api_key",
}

func permanentUpstream(body string) bool {
    b := strings.ToLower(body)
    for _, m := range permanentMarkers {
        if strings.Contains(b, m) {
            return true
        }
    }
    return false
}

The trap inside the fix: the obvious marker to match is the word "quota", and it is the one to leave out. A plain per-minute quota message is genuinely transient, and treating it as permanent turns a two-second blip into a binding that stays cold. Match the phrases that mean money, not the phrases that mean capacity.

Classification, and what catches the case where the body tells you nothing

flowchart LR
  E["Upstream returned
a 429 or a 5xx"] --> B{"Does the body carry
a billing marker?"} B -->|"yes — quota, credit,
hard limit, bad key"| P["Permanent.
Fail over now."] B -->|"no"| T["Transient.
Back off and retry."] T --> N{"Still failing after
the retry ladder?"} N -->|"yes"| BR["Let the breaker
catch the pattern"] N -->|"no"| OK["Serve"]

Body classification cannot catch the third row of that table. A gateway that reports a generic 503 has erased the information you need, and no amount of parsing recovers it. That case is what the next section is for.

Section 03 · the arithmetic

A fallback chain without a breaker charges the full chain to every request.

A fallback chain is usually described as insurance. It is, right up until the moment more than one link is down at once, and then it is a fixed cost added to every single request until a human intervenes.

This is one request, measured, with three of four bindings unable to answer. The chain works exactly as designed. That is the problem.

One turn, walking a chain where only the last link can answer

primary · 503 8.3s
fallback 1 · 400 1.8s
fallback 2 · 429 6.1s
fallback 3 · 200 2.6s

Total 18.8 seconds, of which 16.2 seconds — 86 percent — was spent on providers that could not answer. Repeated on every turn, for every tenant, for 36 hours.

Two separate multipliers are hiding in those bars. The first is the local retry ladder: a status classed as transient is attempted three times with backoff. The second is the gateway's own retry header, which asks it to retry the upstream several times per attempt of ours. A dead binding can therefore absorb nine upstream calls inside what the application thinks is one attempt.

What a breaker changes, and the rule that keeps it safe

A per-binding circuit breaker turns the cost of a dead provider from once per request into once per cooldown window. Ours opens after three consecutive initiation failures, stays open for sixty seconds, and then lets exactly one probe through. A dead binding costs one call a minute instead of one call a turn.

Request path with a per-binding breaker

flowchart TB
  R["Request"] --> C{"Is any binding
still closed?"} C -->|"no — every one is open"| ALL["Ignore the breaker.
Try the whole chain in order."] C -->|"yes"| W["Walk the chain,
skipping open bindings"] W --> T["Attempt binding"] T -->|"success"| RESET["Close the breaker.
Serve the response."] T -->|"initiation failure"| F["Count the failure"] F --> TH{"3 consecutive?"} TH -->|"no"| W TH -->|"yes"| O["Open for 60s.
One probe per window."] O --> W ALL --> T
// The rule that matters most is the last one. A breaker that is allowed
// to skip every candidate converts a provider blip into a total outage —
// strictly worse than the slow chain it replaced.

candidates := append([]binding{primary}, fallbacks...)

useBreaker := false
for _, c := range candidates {
    if !brk.Open(key(c)) {
        useBreaker = true // something is still closed; honour the breaker
        break
    }
}

for _, c := range candidates {
    if useBreaker && !brk.Allow(key(c)) {
        continue // skip the ones we know are dead
    }
    resp, err := try(c)
    if err == nil {
        brk.Success(key(c))
        return resp
    }
    brk.Failure(key(c), err) // context cancellation is not counted
}

The rule worth stealing: when every candidate in the chain is open, the breaker must stand down and let the request try them all anyway. A breaker allowed to skip the entire chain returns a synthetic failure instead of a real one, which is strictly worse than the slow chain it replaced — the user gets nothing instead of something slow, and the logs lose the actual upstream error.

Three implementation details that are easy to get wrong. Key the breaker on the pair of provider and model, not on the provider alone: one gateway credential commonly fronts several models that fail independently, and one dead model should not cool off its healthy siblings. Do not count a cancelled context as a failure, or a user closing a browser tab will trip the breaker on a perfectly healthy provider. And when the half-open probe is handed out, extend the window immediately so that only one caller per cooldown pays for a binding that is still dead.

Section 04 · coverage

The fallback chain existed on one code path out of two.

Streaming chat had a fallback chain, a retry ladder and alerting. The non-streaming completion path returned the primary provider's client and nothing else. It had no fallback at all, and nobody noticed until the primary went down.

This is a structural hazard rather than a bug in the ordinary sense. The two paths exist because streaming and tool-calling have different shapes, and the resilience work was done on the path that user-visible chat traffic takes. Everything else — conversation memory summarisation, scheduled automations, internal tooling, evaluation harnesses — went through the other one and stopped silently the moment the primary failed. They were not erroring loudly; they were returning an error that each caller had been written to degrade past.

Worth auditing in any codebase with more than one way to reach a model: list every entry point into the provider layer, and check each one for the fallback chain, the retry policy, the alerting hook and the cost-attribution context separately. We found the resilience story was complete on one and absent on the other, and the absent one served six distinct internal consumers.

One design note from fixing it. When you wrap a primary client in a fallback chain, a caller that has deliberately pinned a specific model must still get exactly one attempt against that model. An evaluation harness forcing a particular model and silently receiving an answer from a different one produces a score for the wrong thing, which is worse than an error.

Section 05 · verification

The preflight passed, and the model was four times slower.

Before moving the fleet onto a replacement model we built a preflight: send a real request through the candidate binding, confirm it comes back, and confirm it actually calls a tool rather than describing a tool call in prose. It passed. The model was 1.3 to 3.0 seconds. We switched, and reply times went to twenty, thirty-five and sixty-eight seconds.

The preflight was not wrong about anything it measured. It sent a ninety-token prompt, because that is what a health check sends. Production sends the composed system prompt plus retrieved context, which for this agent is around 3,800 tokens. The same model, the same day, the same binding:

One model, two prompt sizes

short probe · 90 input tokens 1.3s
short probe · repeat 3.0s
production · 3.78k input tokens 20.4s
production · 3.79k input tokens 35.1s
production · 3.76k input tokens 67.9s

A 22× spread between the shortest probe and the slowest production turn, on identical infrastructure minutes apart.

The generalisable finding is narrow and worth writing down: latency on a short prompt carries no information about latency on a long one, and the relationship is not a stable multiplier you can calibrate once. A health check that proves a binding is reachable is a useful thing and should not be confused with evidence that it is fast enough to serve.

What we changed: the preflight now sends a production-shaped prompt, and the switch asserts a latency budget rather than only a success. A candidate that answers correctly in thirty-five seconds fails the gate. Reachability, capability and speed are three separate assertions and a green tick on the first two says nothing about the third.

Section 06 · reasoning models

A reasoning model changes the driver contract, not just the answers.

Dropping a reasoning model into a chain built for ordinary chat models breaks in five specific places, and four of them are invisible to the kind of test that sends one short message and checks for a reply.

It refuses an explicit temperature

HTTP 400 · invalid temperature

The model accepts only the vendor default. Every request that sets a temperature fails, and a bare chat probe that sets none succeeds — so the model looks healthy in testing and broken in production. We had nineteen call sites setting a temperature.

Fix Normalise per model family in the driver, the same way an output-cap field is already normalised. Drop the field rather than passing a value through.

Thinking consumes the output budget before any text exists

empty content · finish_reason: length

With a small output cap the model returned an empty string and a length finish reason. The budget had gone entirely on reasoning tokens. On an open-ended question it spent 470 of 555 completion tokens thinking, which is 85 percent of what you are billed for and none of what you display.

Fix Raise every output cap when a reasoning model is in the chain, and treat an empty completion as its own alarm rather than a successful turn.

Reasoning arrives on the stream as its own delta type

29 of 40 chunks were reasoning

A short greeting streamed as 40 chunks: 29 carried a reasoning field and 9 carried content. A parser that reads only the content field is safe and silently correct. A parser that concatenates every delta leaks the chain of thought to the end user.

Fix Read the content field by name. Never concatenate an unknown delta shape.

Time to first byte is the whole latency

88–95% of each measurement

Because the model thinks before it emits, streaming buys nothing perceptually. The user watches a typing indicator for the entire reasoning phase and then receives the answer almost at once.

Fix Budget for total latency, not for first token. Streaming is not a mitigation here.

Turning thinking off is not a reliable lever

3.4s once · then 12.4s, 9.8s, 20.5s

A disable flag produced one fast result that would not reproduce across the next three calls, for replies of only 71 to 95 output tokens. The bottleneck was serving capacity, not reasoning volume.

Fix Measure the flag across a run of calls before believing it. One fast sample is a cache hit, not a finding.

The empty completion is the one that will catch you

A successful 200 with an empty content field and a length finish reason is not a shape most chat code is written to expect. It is what a reasoning model returns whenever the output cap is smaller than the thinking it wanted to do.

// A reasoning model with a small output cap returns this. It is a
// successful HTTP 200 with nothing to show the user.

"choices": [ ... "message": ... "content": "" ... ]
"finish_reason": "length"
"usage": ... "completion_tokens": 20,
              "completion_tokens_details": ... "reasoning_tokens": 17

Any pipeline that treats "no error" as "we have a reply" will persist an empty message, increment its success counter, and record a healthy latency. The general form of that lesson keeps recurring in this stack: a failure path that increments a success metric removes itself from the list of suspects and outlives several rounds of debugging.

The parameter that is refused, not ignored

Providers vary in whether an unsupported parameter is dropped politely or rejected with a 400, and the rejecting ones are far more dangerous because the failure is conditional on the caller. Our replacement model rejected any explicit temperature. Nineteen call sites in the codebase set one, and the health check set none — so the model passed every test we had and failed every path that mattered. Normalising this per model family in the driver, the same way an output-cap field name is already normalised, is the only place the fix belongs. Doing it in nineteen call sites is nineteen chances to miss one.

Section 07 · behaviour

Three families, one strict prompt, three different failures.

The test agent is a deliberately demanding case: a knowledge desk whose system prompt requires it to name the source document behind every claim, end with one of four literal confidence words, keep replies to two to four sentences, and say plainly when the record does not cover something. Every model held some of that and broke a different part of it.

Family A
Large proprietary generalist
Best rule adherence
latency1.2 s avg
output6–161 tok
prompt cache~50% of input
Family B
Mid-size open-weight
Fast, confidently wrong
latency1.9–7.9 s
costlowest by 10×
usage framenone returned
Family C
Reasoning model
Most honest, unusably slow
latency20–68 s
words/sentence26.5–41.3
reasoningbilled as output

Family A · it obeys the rule you wrote, not the rule you meant

The large proprietary generalist was the best-behaved of the three and the only one that never invented an answer. It also failed on its very first run, by paraphrasing the confidence vocabulary and appending exactly the filler phrase the prompt banned. The rule as originally written described an intent. Rewriting it to list the four literal tokens, pin their position in the reply, and name the banned alternative fixed it permanently.

It was also the only model of the three to find a conditional carve-out in the formatting rule. The prompt forbids bulleted lists unless the answer is genuinely a sequence of steps someone performs in order, capped at six. Asked for a procedure, this model produced a numbered list; asked anything else, prose. Neither other model ever triggered the exception, which means a conditional rule is effectively invisible to them and the behaviour you get is whichever branch is stated first.

Cost note worth planning around: roughly half the input tokens on each turn came back marked as cache hits once the system prompt stabilised, at a fraction of the uncached rate. A prompt that is regenerated or reordered on every turn forfeits that discount silently. Prompt stability is a cost lever, not just a quality one.

Family B · perfect shape, invented content

The mid-size open-weight model kept the length rule better than the large model did: short sentences, no headings, no filler. Shape was never the problem. The problem was that the shape stayed intact while the content became fiction.

Asked who owned a particular internal procedure document, it answered with a person's name, cited the correct document, gave the correct review date, dropped the caveat that the document carries about its own status, and tagged the whole thing with the highest confidence value in the vocabulary. The document's header names a role, not a person. The name it supplied appears elsewhere in the same prompt, as a contact. The model reached for the nearest human-shaped token and formatted it impeccably.

This is the worst possible output for a sourcing agent. A refusal is safe. A hedge is readable. A well-formatted, correctly-cited, confidently-tagged wrong answer defeats every control the prompt installs, because every signal a reader uses to judge it is present and correct. The large model answered the same question correctly from the same corpus.

Two further behaviours from the same family, both reproducible. It re-introduces itself when retrieval comes back thin — with nothing retrieved to say, it falls back on the most salient thing in the prompt, which is the identity block at the top. And it answers when it should ask: a user typed a single word with no intent in it and received a full policy statement in reply, assembled from whatever was in the context window.

Its confidence tags were syntactically perfect and statistically meaningless. The same question, against the same documents, produced two different confidence values on two different days. The control was present in every reply and measuring nothing.

Family C · the best epistemics, at forty seconds a turn

The reasoning model produced the single best behaviour we saw from any of the three, unprompted. Asked about a document whose text had not come back in the retrieved passages, it distinguished between the record does not contain this and my search did not return it, said which one it was facing, and offered three alternative search terms. Nothing in the prompt asks for that distinction. It is exactly what a knowledge desk should do and the one behaviour we have not found a way to teach reliably. The mid-size model, on the same question, invented an owner instead.

It also comprehensively gamed the length rule. Asked a one-line question, it produced three sentences — technically compliant — running to 124 words at 41.3 words per sentence, and 2,385 output tokens. On top of that it re-answered the previous turn before addressing the current one, opening with an explicit "catching up in order". That is reasoning-model behaviour showing through: it reconsiders the whole conversation rather than the last message, and conversational state you thought was settled gets relitigated.

The matrix

Behaviour under one promptA · proprietaryB · open-weightC · reasoning
Emits the literal confidence token after a rewrite yes, uncalibrated yes
Holds the sentence limit yes yes by the letter only
Refuses when the record is silent yes invents an answer yes, and says why
Separates retrieval gap from absence no no yes
Keeps the persona out of the answer yes leaks on thin recall yes
Carries a caveat the document states yes drops it n/a
Honours a conditional formatting rule yes never triggered it never triggered it
Latency on a production prompt 1.2 s 2–4 s 20–68 s
Cost observable in the ledger yes no usage frame yes

The observation underneath the table: speed, rule-following and honesty were distributed across three different models, and no single one had all three. If your agent's value proposition is that it tells you where an answer came from, a model that is fast and cheap and invents attributions is not a cheaper version of the right model. It is a different product.

Section 08 · prompt authoring

Writing rules that survive a model swap.

Every prompt in production is tuned against a model, whether or not anyone intended that. These are the four properties that made a rule portable in our testing, and the one that made a rule fragile.

State the vocabulary, not the intent

The single highest-leverage rewrite we made. A rule describing what you want produces a paraphrase; a rule listing the exact tokens produces the tokens.

# What did not work — an instruction stated as intent.
  Be clear about how confident you are in each answer.

# What worked — the same rule stated as vocabulary, with the
# position in the reply pinned, and a banned alternative named.
  End any answer containing a fact with a confidence tag on its own,
  as the LITERAL last words, in exactly this form:
    Confidence: certain | mostly | thin | guessing
  Those four words are the whole vocabulary. Do not substitute a
  paraphrase. Writing "this is confident" or "I'm fairly sure" is
  wrong even when the meaning is the same, because the point is
  that a reader can scan for the word.

The clause that did the most work is the last one — naming the failure explicitly, and saying why the literal form matters. A rule that explains its own purpose survives a model that is trying to be helpful.

Anything countable will be optimised, so bound it twice

A length rule stated in sentences is satisfied by making the sentences longer. This is not the model being adversarial; it is the model optimising the only quantity the rule names.

# Stated in one unit — a reasoning model satisfies it at four times
# the length, because sentences are the only thing being counted.
  A normal reply is 2 to 4 SENTENCES.

# Stated in two units — now the rule constrains the thing you
# actually cared about.
  A normal reply is 2 to 4 SENTENCES and under 80 words.
  Count both before you send.

The same shape applies elsewhere. A rule capped in steps gets longer steps. A rule capped in paragraphs gets denser paragraphs. Name the unit you care about, then name the one the model would otherwise expand into.

A negative example moves behaviour further than an instruction

When a rule was being ignored, the change that fixed it was not stronger wording. It was pasting the actual over-long output verbatim into the prompt, labelled as the thing not to do, beside a corrected version of the same answer. An abstract instruction competes with the model's prior; a concrete pair of the same content in two shapes does not.

Position encodes precedence

Our formatting rule sits last in the composed prompt, after the persona, specifically so that it outranks the persona when the two conflict. That ordering held for two of the three families. For the third it made no observable difference. Position is a real lever and not a reliable one, which is an argument for stating precedence in words as well as in order.

The fragile property: a conditional rule — do X unless Y, in which case do Z — was honoured correctly by exactly one of three models. The other two never triggered the exception at all. If a conditional matters, test that the exception branch actually fires on the model you are running, because the silent behaviour is that only the first branch exists.

Section 09 · what we changed

The checklist we wish we had run first.

01

Classify upstream failures on the body, not the status

Billing exhaustion arrives as a 429 from one vendor and a 400 from another. Keep the permanent-marker list narrow enough that an ordinary rate limit is still retried.

02

Put a circuit breaker in front of every fallback chain

Key it on provider and model together. Do not count cancelled contexts. Make it stand down entirely when every candidate is open, so it can never black-hole the chain.

03

Audit every entry point into the model layer, not just the busy one

Fallbacks, retries, alerting and cost attribution each need checking per path. Ours was complete on the streaming path and entirely absent on the non-streaming one.

04

Preflight with a production-shaped prompt and a latency budget

Reachability, tool-calling and speed are three assertions. A ninety-token probe answers only the first two, and a 22× latency spread is hiding behind them.

05

Normalise per-model parameter quirks in the driver

Refused temperatures, renamed output caps and reasoning-token budgets belong in one place. A quirk handled at the call sites is one missed call site away from a conditional outage.

06

Treat an empty completion as an alarm, not a turn

A 200 with no content and a length finish reason is a reasoning model running out of budget before it says anything. Any code that reads "no error" as "we have a reply" will record it as a success.

07

Price every model in the chain, including the free-looking one

Spend reading zero for three days looked like a quiet week. It was an empty balance on one model and an unpriced fallback on the next. A fallback nobody has costed is a fallback nobody is monitoring.

08

Re-run your prompt regression suite on the fallback, not only on the primary

The fallback is the model your users get on your worst day. Ours had never been evaluated against the prompt it ended up serving, and it was the one that invented attributions.

The last one is the one we would put first if we were starting again. A fallback chain is a list of models your users will eventually talk to. Every one of them deserves the same evaluation as the primary, because the day it serves traffic is by definition the day nobody is watching closely.

Related reading

The per-turn model, latency and cost figures these findings were drawn from come from the pipeline described in Conversation Observability. The driver-and-factory shape that made the per-model parameter normalisation a one-file change is written up for messaging in the Common Channel Wrapper. A different latency investigation, on the voice side, is in From a silent call to 820 milliseconds.

// field notes · model routing · figures from production traffic

// edit, open a PR, ships on merge to main