Skip to main content
Omazy Engineering

Context · caching vs compression

Two levers for a growing brief.

The agent brief is re-sent on every turn, so its size is a bill paid over and over. The safe compaction saves almost nothing on a clean brief. The two levers that matter are provider prompt caching, which cuts cost at zero behavioural risk, and reviewed compression, which cuts size but must never reword a rule in secret.

// prompt caching · reviewed compression · invariant guard // RFC 2222 · Track A · shipped on main · deployed

Section 01 · the problem

A brief you re-send every turn is a bill you pay every turn.

The system brief that gives an agent its identity, rules, and worked examples is composed once and then sent to the model on every single turn. As a brief grows past a few thousand tokens, that is a recurring per-turn cost, a per-turn prefill latency, and a standing draw against the model's usable context window.

The obvious lever is the wrong one. We shipped a provably lossless whitespace compaction (Compact, in internal/agentbrief/compact.go) and measured what it saved on a real, mature brief: zero tokens. The composer already emits tight markdown, so there is no free whitespace to reclaim. The tokens that could shrink are the prose in Instructions and Samples, which are exactly the tokens where a reword changes behaviour. That honest result is what forces the real question: if the safe lever saves nothing, what actually moves the number without moving the agent?

Section 02 · two levers, not one

Caching cuts the cost. Compression cuts the size. They are not the same tool.

These are complementary, not competing. Caching makes the same prefix cheap to re-send; compression makes the prefix smaller. One is a runtime property with zero behavioural risk; the other is an authoring decision that has to be reviewed.

flowchart LR
  subgraph AUTHOR["① AUTHORING TIME · operator"]
    I["brief items
(agent_prompt_items)"] --> C["compose"] C --> K["Compact · lossless
whitespace only"] C -.->|"opt-in per prompt"| Z["Compress · model
reviewed diff"] K --> P["publish → agents.system_prompt"] Z -.->|"accept = Update"| P end subgraph TURN["② EVERY CUSTOMER TURN · runtime"] P --> PRE["stable system prefix"] PRE --> PROV["provider via CF AI Gateway /compat"] PROV -->|"cache hit?"| BR[("bot_responses
cached_input_tokens")] end
Fig 1. Both levers act away from the customer turn. Compaction and compression happen at authoring time; caching is a transparent property of the runtime call. The customer reply path reads an already-composed, already-compressed static prompt and does zero extra work.
LeverFixesBehavioural riskWhen it runs
Prompt cachingCostNoneTransparent, every turn
Whitespace compactionNothing on a clean briefNone (lossless)Publish
Model compressionSize (cost, latency, context budget)Real: a reworded rule is a changed ruleAuthoring time, operator-triggered

Section 03 · make it measurable first

You cannot tune what you cannot see. So we made cache hits visible.

The brief is the textbook case for provider prompt caching: a long, stable prefix re-sent constantly. On a cache hit, cached input is billed at roughly a tenth of the rate. But we had never confirmed whether cache hits actually flow through the path we route on, because the number was being thrown away.

Production routes every model through the Cloudflare AI Gateway /compat endpoint in OpenAI-compatible wire format. The usage frame on that path carries the cache-served portion of the input as prompt_tokens_details.cached_tokens (the Anthropic-native path reports the same thing as cache_read_input_tokens). The stream parser read prompt_tokens and completion_tokens and dropped the cached breakdown on the floor, so a cache hit was invisible.

The change is deliberately small and additive: parse the cached count already on the wire, thread it through the dispatcher, and persist it in a new ClickHouse column bot_responses.cached_input_tokens (migration 011, auto-applied on boot). It is a subset of input tokens, not additive. Nothing about what is sent upstream changes. The whole spike then reduces to a one-line query: read the column, and a non-zero value on a repeat turn is proof caching is flowing.

The principle: the first move on a performance question is not to optimise, it is to instrument. A cache lever with no cache metric is a guess. The cheapest, zero-risk change was to capture a number that was already being sent and then discarded.

Section 04 · compression, reviewed

The model may propose a shorter brief. It may never apply one.

Model-driven compression is the only lever that shrinks the prose caching cannot. It is also the dangerous one, so the entire flow is built around a single rule: nothing reaches the live agent without an explicit human accept.

flowchart TB
  REQ["POST /agent/brief/items/:id/compress"] --> RUN["Service.Compress
runs CompressSystem, writes nothing"] RUN --> G{"deterministic invariant guard
numbers · prices · urls · code fences"} G -->|"drops → flags"| DIFF["operator reviews the diff"] G -->|"clean"| DIFF DIFF -->|"Accept"| U["Update → auto-version
earlier body one Restore away"] DIFF -->|"Discard"| N["nothing written"] U --> LIVE["live agent"]
Fig 2. Service.Compress runs the meta-prompt and writes nothing. The operator reviews a diff; Accept goes through the ordinary item Update, which auto-versions, so the pre-compression body is one Restore away. Discard leaves the brief untouched.

Off the hot path, on purpose

Compression is an authoring action, triggered by an operator in the console, never automatically and never on a customer turn. By the time a customer chats, the brief is already compressed and static. The only wait is the operator's, for a few seconds, when they request a proposal.

Undo is free because it is not special

Accepting a proposal is just a normal item update. The brief already snapshots the prior version on every content change and keeps three. So the safety net for compression is the same keep-three history every edit already uses, no new mechanism.

Section 05 · the deterministic guard

A human skimming a diff misses a dropped price. Code does not.

A reworded paragraph reads fine even when it quietly lost a number. So before the operator ever sees a proposal, a deterministic guard runs, with no model involved, and flags any fact the shorter version dropped.

The guard extracts every fenced code block, URL, price, and number from the original, then checks each survives in the compressed output. It works by membership, not by count: a value that appears at least once in the shorter version is not flagged, so ordinary rewording stays quiet and only a genuine drop is loud. Code, URLs, and prices are scanned first and their spans blanked before the loose-number pass, so a digit inside a surviving price is never double-reported. It is cheap, needs no LLM, and catches the single highest-risk silent change: a lost link or figure.

Where the line is drawn (decision D2): rule-type prompts are held out of model compression entirely in v1. A rule is a limit the agent must never cross, and it is where a reword is most dangerous, so a rule is whitespace-compacted only and its wording is left byte-for-byte. Every other type is eligible, and every proposal, even a guard-clean one, is still reviewed. There is no auto-accept, because auto-accept would reintroduce the exact silent-drift risk the feature exists to prevent.

Section 06 · why both, given the growth

Size hits three axes. Caching only fixes one of them.

It is tempting to pick one lever. The growth of the brief is exactly why you cannot.

Axis a large brief hurtsCachingCompression
Per-turn costFixes it (cached input billed ~10x cheaper)Also helps (fewer tokens)
Prefill latencyFixes it only on a hit (cold cache still pays)Fixes it always (shorter prefix)
Context-window budgetDoes not fix it (a cached prefix still occupies the window)Only this fixes it

Do caching first, because it lands the cost half of the problem with zero behavioural risk. Do compression for the two axes caching cannot touch: guaranteed latency and the context budget. The bigger the brief gets, the stronger the compression case becomes on exactly those two.

Section 07 · design principles

Two rules the whole design obeys.

Nothing expensive on the customer turn

Compression runs at authoring time; caching is transparent. The reply path reads a static, already-processed prompt. A shorter brief is marginally faster to prefill every turn, and compression never adds a millisecond to a reply.

A reworded rule is a changed rule

Compression proposes, it never applies. The operator reviews a diff, the guard flags dropped facts, rules are excluded, and accept is a versioned update. Every path assumes the shorter text could be subtly wrong until a human confirms it is not.

Section 08 · implementation surface

Where it lives.

Caching instrumentation is in the LLM drivers and the dispatcher; compression is the agent-brief service and one console flow. One ClickHouse column, no Postgres migration.

Caching (measurement)

  • internal/llm/driver_cf_aig.go: parse prompt_tokens_details.cached_tokens off the usage frame
  • internal/llm/driver.go, tools.go: CachedInputTokens on Chunk and Usage
  • cmd/server/wiring/ouchat_dispatcher.go: thread the cached count into the fact write
  • db/clickhouse/migrations/011_bot_responses_cached_tokens.sql: the column, applied on boot

Compression (reviewed)

  • internal/agentbrief/compress.go: Service.Compress + the invariant guard
  • internal/agentbrief/prompts.go: the CompressSystem meta-prompt
  • internal/workspaceapps: the POST /agent/brief/items/:id/compress route
  • brand/agent/prompts/_components/brief-manager.tsx: the Shorten action and review modal

Sources & further reading

The lossless and lossy halves each carry their reasoning in-file; the RFC records the decisions and the honest measurements.

  • docs/rfc2222-agent-context-economics-and-conversation-observability.md: the plan, decisions, and deploy sequence
  • internal/agentbrief/compact.go: the lossless half, with the "compaction saved 0" result stated plainly
  • GitHub #142 (Agent Brief), #154 (compression)

Related engineering reading: Conversation Observability (Track B, where cached_input_tokens surfaces as operator cost), and the Dashboard & Telemetry Engine (the bot_responses fact table this writes to).