Context · caching vs compression
Two levers for a growing brief.
The agent brief is re-sent on every turn, so its size is a bill paid over and over. The safe compaction saves almost nothing on a clean brief. The two levers that matter are provider prompt caching, which cuts cost at zero behavioural risk, and reviewed compression, which cuts size but must never reword a rule in secret.
Section 01 · the problem
A brief you re-send every turn is a bill you pay every turn.
The system brief that gives an agent its identity, rules, and worked examples is composed once and then sent to the model on every single turn. As a brief grows past a few thousand tokens, that is a recurring per-turn cost, a per-turn prefill latency, and a standing draw against the model's usable context window.
The obvious lever is the wrong one. We shipped a provably lossless whitespace compaction (Compact, in internal/agentbrief/compact.go) and measured what it saved on a real, mature brief: zero tokens. The composer already emits tight markdown, so there is no free whitespace to reclaim. The tokens that could shrink are the prose in Instructions and Samples, which are exactly the tokens where a reword changes behaviour. That honest result is what forces the real question: if the safe lever saves nothing, what actually moves the number without moving the agent?
Section 02 · two levers, not one
Caching cuts the cost. Compression cuts the size. They are not the same tool.
These are complementary, not competing. Caching makes the same prefix cheap to re-send; compression makes the prefix smaller. One is a runtime property with zero behavioural risk; the other is an authoring decision that has to be reviewed.
flowchart LR
subgraph AUTHOR["① AUTHORING TIME · operator"]
I["brief items
(agent_prompt_items)"] --> C["compose"]
C --> K["Compact · lossless
whitespace only"]
C -.->|"opt-in per prompt"| Z["Compress · model
reviewed diff"]
K --> P["publish → agents.system_prompt"]
Z -.->|"accept = Update"| P
end
subgraph TURN["② EVERY CUSTOMER TURN · runtime"]
P --> PRE["stable system prefix"]
PRE --> PROV["provider via CF AI Gateway /compat"]
PROV -->|"cache hit?"| BR[("bot_responses
cached_input_tokens")]
end
| Lever | Fixes | Behavioural risk | When it runs |
|---|---|---|---|
| Prompt caching | Cost | None | Transparent, every turn |
| Whitespace compaction | Nothing on a clean brief | None (lossless) | Publish |
| Model compression | Size (cost, latency, context budget) | Real: a reworded rule is a changed rule | Authoring time, operator-triggered |
Section 03 · make it measurable first
You cannot tune what you cannot see. So we made cache hits visible.
The brief is the textbook case for provider prompt caching: a long, stable prefix re-sent constantly. On a cache hit, cached input is billed at roughly a tenth of the rate. But we had never confirmed whether cache hits actually flow through the path we route on, because the number was being thrown away.
Production routes every model through the Cloudflare AI Gateway /compat endpoint in OpenAI-compatible wire format. The usage frame on that path carries the cache-served portion of the input as prompt_tokens_details.cached_tokens (the Anthropic-native path reports the same thing as cache_read_input_tokens). The stream parser read prompt_tokens and completion_tokens and dropped the cached breakdown on the floor, so a cache hit was invisible.
The change is deliberately small and additive: parse the cached count already on the wire, thread it through the dispatcher, and persist it in a new ClickHouse column bot_responses.cached_input_tokens (migration 011, auto-applied on boot). It is a subset of input tokens, not additive. Nothing about what is sent upstream changes. The whole spike then reduces to a one-line query: read the column, and a non-zero value on a repeat turn is proof caching is flowing.
The principle: the first move on a performance question is not to optimise, it is to instrument. A cache lever with no cache metric is a guess. The cheapest, zero-risk change was to capture a number that was already being sent and then discarded.
Section 04 · compression, reviewed
The model may propose a shorter brief. It may never apply one.
Model-driven compression is the only lever that shrinks the prose caching cannot. It is also the dangerous one, so the entire flow is built around a single rule: nothing reaches the live agent without an explicit human accept.
flowchart TB REQ["POST /agent/brief/items/:id/compress"] --> RUN["Service.Compress
runs CompressSystem, writes nothing"] RUN --> G{"deterministic invariant guard
numbers · prices · urls · code fences"} G -->|"drops → flags"| DIFF["operator reviews the diff"] G -->|"clean"| DIFF DIFF -->|"Accept"| U["Update → auto-version
earlier body one Restore away"] DIFF -->|"Discard"| N["nothing written"] U --> LIVE["live agent"]
Service.Compress runs the meta-prompt and writes nothing. The operator reviews a diff; Accept goes through the ordinary item Update, which auto-versions, so the pre-compression body is one Restore away. Discard leaves the brief untouched.Off the hot path, on purpose
Compression is an authoring action, triggered by an operator in the console, never automatically and never on a customer turn. By the time a customer chats, the brief is already compressed and static. The only wait is the operator's, for a few seconds, when they request a proposal.
Undo is free because it is not special
Accepting a proposal is just a normal item update. The brief already snapshots the prior version on every content change and keeps three. So the safety net for compression is the same keep-three history every edit already uses, no new mechanism.
Section 05 · the deterministic guard
A human skimming a diff misses a dropped price. Code does not.
A reworded paragraph reads fine even when it quietly lost a number. So before the operator ever sees a proposal, a deterministic guard runs, with no model involved, and flags any fact the shorter version dropped.
The guard extracts every fenced code block, URL, price, and number from the original, then checks each survives in the compressed output. It works by membership, not by count: a value that appears at least once in the shorter version is not flagged, so ordinary rewording stays quiet and only a genuine drop is loud. Code, URLs, and prices are scanned first and their spans blanked before the loose-number pass, so a digit inside a surviving price is never double-reported. It is cheap, needs no LLM, and catches the single highest-risk silent change: a lost link or figure.
Where the line is drawn (decision D2): rule-type prompts are held out of model compression entirely in v1. A rule is a limit the agent must never cross, and it is where a reword is most dangerous, so a rule is whitespace-compacted only and its wording is left byte-for-byte. Every other type is eligible, and every proposal, even a guard-clean one, is still reviewed. There is no auto-accept, because auto-accept would reintroduce the exact silent-drift risk the feature exists to prevent.
Section 06 · why both, given the growth
Size hits three axes. Caching only fixes one of them.
It is tempting to pick one lever. The growth of the brief is exactly why you cannot.
| Axis a large brief hurts | Caching | Compression |
|---|---|---|
| Per-turn cost | Fixes it (cached input billed ~10x cheaper) | Also helps (fewer tokens) |
| Prefill latency | Fixes it only on a hit (cold cache still pays) | Fixes it always (shorter prefix) |
| Context-window budget | Does not fix it (a cached prefix still occupies the window) | Only this fixes it |
Do caching first, because it lands the cost half of the problem with zero behavioural risk. Do compression for the two axes caching cannot touch: guaranteed latency and the context budget. The bigger the brief gets, the stronger the compression case becomes on exactly those two.
Section 07 · design principles
Two rules the whole design obeys.
Nothing expensive on the customer turn
Compression runs at authoring time; caching is transparent. The reply path reads a static, already-processed prompt. A shorter brief is marginally faster to prefill every turn, and compression never adds a millisecond to a reply.
A reworded rule is a changed rule
Compression proposes, it never applies. The operator reviews a diff, the guard flags dropped facts, rules are excluded, and accept is a versioned update. Every path assumes the shorter text could be subtly wrong until a human confirms it is not.
Section 08 · implementation surface
Where it lives.
Caching instrumentation is in the LLM drivers and the dispatcher; compression is the agent-brief service and one console flow. One ClickHouse column, no Postgres migration.
Caching (measurement)
internal/llm/driver_cf_aig.go: parseprompt_tokens_details.cached_tokensoff the usage frameinternal/llm/driver.go,tools.go:CachedInputTokenson Chunk and Usagecmd/server/wiring/ouchat_dispatcher.go: thread the cached count into the fact writedb/clickhouse/migrations/011_bot_responses_cached_tokens.sql: the column, applied on boot
Compression (reviewed)
internal/agentbrief/compress.go:Service.Compress+ the invariant guardinternal/agentbrief/prompts.go: theCompressSystemmeta-promptinternal/workspaceapps: thePOST /agent/brief/items/:id/compressroutebrand/agent/prompts/_components/brief-manager.tsx: the Shorten action and review modal
Sources & further reading
The lossless and lossy halves each carry their reasoning in-file; the RFC records the decisions and the honest measurements.
docs/rfc2222-agent-context-economics-and-conversation-observability.md: the plan, decisions, and deploy sequenceinternal/agentbrief/compact.go: the lossless half, with the "compaction saved 0" result stated plainly- GitHub #142 (Agent Brief), #154 (compression)
Related engineering reading: Conversation Observability (Track B, where cached_input_tokens surfaces as operator cost), and the Dashboard & Telemetry Engine (the bot_responses fact table this writes to).