Skip to main content
Omazy Engineering

Observability · customer trace + operator provenance

One conversation, two audiences.

A customer needs a way to report a bad answer that a human can resolve. An operator needs to see which model answered, how long it took, and what it cost. The same conversation owes each of them something different, so the surface is split: a copyable reference for the customer, full per-turn provenance for the operator, and never the two mixed.

// trace code · reply provenance · bot_responses · llmprice // RFC 2222 · Track B · shipped on main · deployed

Section 01 · the problem

The thread has to stay clean, and still be debuggable.

The widget conversation is deliberately minimal, and it must stay that way. But two people need more from it than it shows. A customer who gets a wrong answer has no way to report it that a human can act on. An operator has no per-turn view of which model answered, how long it took, or what it cost.

The wrong fix is to pour metadata into the thread. Model names and token counts mean nothing to a customer and clutter the one surface that has to stay simple. The right fix starts from a different question: who is this information for? A customer needs a reference they can quote. An operator needs the numbers. Those are different people looking at different screens, so they get different surfaces.

Section 02 · audience decides the surface

Split it by who is looking.

The same AI reply owes the customer a way to report and the operator a way to account. Neither should ever see the other's view.

flowchart TB
  TURN["one AI reply"] --> CUST
  TURN --> OP
  subgraph CUST["CUSTOMER · the widget thread (sacred, minimal)"]
    TD["thumbs-down"] --> REF["Ref 3DE8-C322
copyable, resolves to session + message"] end subgraph OP["OPERATOR · the console inbox"] PROV["gpt-4o · 1.2s · 4.0k→120 · $0.0013"] end REF -.->|"customer quotes the code"| PROV
Fig 1. The customer surface carries a trace reference and nothing more. The operator surface carries the full per-turn picture. The only link between them is the customer quoting the code, which an operator resolves straight to the exact session and message.

The thread is sacred

Nothing new appears in the conversation until the customer chooses to report. The trace affordance is tucked inside the feedback control they already reach for, so the default view is exactly as clean as it was before.

Numbers are operator data

Token cost, model, and latency live on operator surfaces only. The customer sees a reference that means something to support, never figures that mean nothing to them.

Section 03 · the customer side

A code the customer can quote. Nothing more.

When a customer marks a reply unhelpful, a single tappable pill appears: a short reference like Ref 3DE8-C322 that copies a fuller, resolvable string to the clipboard. That is the entire feature on the customer side.

The code is derived on the client from ids already on the message, the session prefix and the message suffix, so there is no server round-trip and no new table. It carries no personal data. When the customer follows up through any channel and quotes the code, an operator resolves it back to the exact session and message. It appears only after a thumbs-down, so the thread stays clean until the customer decides to report, and it shows no model, tokens, or cost, because none of that helps a customer and all of it would clutter the bubble.

The principle: the smallest thing that makes a bad answer reportable is a resolvable reference, not a form and not a metadata dump. Start code-only. A persisted report inbox can be added later if operators want one; it was deliberately left out of v1 so the customer change is pure client code that ships on its own.

Section 04 · the operator side

Model, latency, tokens, cost. Under every AI reply.

In the console inbox, each AI reply gets a compact line beneath it: the model that answered, how long it took, tokens in and out, any cache-served tokens, and the list-price cost of the turn. It is almost entirely a display task over data already recorded.

flowchart LR
  T["transcript · AI message ids"] --> Q["GET /inbox/:session/provenance?ids=…"]
  Q --> CH[("bot_responses
WHERE message_id IN (…)")] CH --> PR["price with internal/llmprice"] PR --> LINE["provenance line under the AI bubble"]
Fig 2. The console sends the AI message ids in view to a provenance endpoint, which reads the bot_responses fact table by message_id and prices each row with internal/llmprice. Only AI turns have a row, so ids with no fact are simply absent from the response.

The join is exact: bot_responses.message_id is the ULID of the reply message, so there is no fuzzy matching. The endpoint reads only the ids the operator has on screen, a bounded set, and prices them with the same per-model rate table the usage dashboards use, so a turn's cost here agrees with the monthly rollup. Pricing is at list rate today, with no markup; a model without a published rate is flagged as an estimate rather than presented as a fact.

Section 05 · reuse, not new capture

Almost none of this is new data.

The whole operator track is a read over facts the platform already writes. Nothing new is captured on the serving path.

The facts already exist

bot_responses already records model, response time, and input/output tokens per AI turn. The pricing table already turns tokens into cost for the usage dashboards. The provenance view is a new read endpoint and a new line of UI, not a new pipeline.

Cache visibility comes along

The cache-served token count added in Track A rides in the same fact row, so the operator line shows "N cached" when a turn hit the provider cache. The two tracks share one fact table and reinforce each other.

Section 06 · failure modes, by design

Telemetry degrades quietly. It never breaks the inbox.

Best-effort read

If the analytics store is not wired, or the fact has not flushed yet, the endpoint returns an empty map, not an error. The inbox shows "no provenance" rather than a broken pane. A missing metric never blocks an operator from reading a conversation.

Scoped and bounded

Provenance sits behind the same per-app ownership check as the transcript, so one app's facts never leak into another's inbox, and the id list is capped so a caller cannot force an unbounded scan over the fact table.

The customer report is code-only by the same instinct: no new table means the widget change carries no migration and no server dependency, so it ships as pure client code and cannot half-break a backend it does not touch.

Section 07 · implementation surface

Where it lives.

The customer trace is one widget component; the operator provenance is a read endpoint plus one line of console UI. No migration on either side.

Customer (widget)

  • widget/src/ui/message-row.tsx: the trace pill, code derivation, clipboard copy
  • widget/src/i18n.ts: the report strings, fully localisable
  • Client-only: no endpoint, no table, no round-trip

Operator (middleware + console)

  • internal/inbox/provenance.go: the ProvenanceReader + the read handler
  • cmd/server/wiring/inbox.go: the ClickHouse-backed reader, priced with llmprice
  • lib/api/conversations.ts, inbox/_hooks/use-provenance.ts: fetch by message id
  • inbox/_components/message-bubble.tsx: the compact provenance line

Sources & further reading

The audience split and the reuse-only stance are recorded in the RFC; the reader carries its ownership and pricing rules in-file.

  • docs/rfc2222-agent-context-economics-and-conversation-observability.md: Track B, decisions D3 (report persists vs code-only)
  • internal/inbox/provenance.go: the read model, the message-id join, the best-effort contract
  • GitHub #133 / #138: conversation intelligence (intent, topic, summary), which renders alongside provenance when it lands

Related engineering reading: Agent Context Economics (Track A, where the cached_input_tokens shown here is captured), and the Dashboard & Telemetry Engine (the bot_responses fact table and the cost-tier model behind the pricing).