RFC 2212 · Telemetry & dashboards
Measure everything. Slow nothing.
How the workspace dashboards collect, store and serve every metric behind a hard boundary from the AI core — where the cost of producing a stat decides who can see it, and shipping a feature can never break what is live.
Section 01 · the problem
A dashboard that lags reality is useless. One that slows the AI is worse.
Business owners want to watch their business live — conversations, revenue, ROI. But the AI serving path is sacred: a chat reply must never wait on a metrics write, and a heavy dashboard query must never lock a table the agent loop needs.
So the whole system is built around one boundary. The AI core does exactly one measurement thing — a sub-millisecond counter bump — and everything else happens later, on separate workers, against a separate database. On top of that we layer a second rule the industry usually forgets: producing a stat costs money, so what it costs decides who can see it. Real-time is opt-in and tier-locked, never a flat tax on every request.
Section 02 · architecture
Three planes, one hard boundary.
Emit is the only thing on the request path, and it is O(1). Process and Serve run entirely off it — a metrics backlog or a ClickHouse outage is invisible to a customer mid-chat.
flowchart LR
subgraph HOT["① EMIT · request path · O(1)"]
AI["AI core / chat reply"] -->|"Redis INCR (cached snapshot)"| R[("Redis counters")]
AI -->|"append"| B["in-proc buffer"]
WDG["Widget / website"] -->|"sendBeacon"| COL["edge collector · Worker
(separate from origin CPU)"]
end
subgraph PROC["② PROCESS · isolated workers (Asynq low)"]
B -->|"batch 200/5s"| W["buffered writers"]
COL --> W
W -->|"on CH failure"| OB[("analytics_outbox · Postgres")]
OB -->|"drain @30s · SKIP LOCKED"| CH[("ClickHouse · omazy_analytics")]
W --> CH
CH --> MV[("materialized views · rollups")]
end
subgraph SERVE["③ SERVE · read-only"]
MV --> API["/metrics API · ms point-reads"]
R --> API
API --> D["dashboard widgets"]
MV -.->|"opt-in · ref-counted"| PUB["realtime publisher (C3)"]
PUB -.-> D
end
Why chat can't be starved
Chat/AI is synchronous HTTP/SSE — it never enters the job queue. All analytics runs on Asynq's weighted low queue (internal/jobs/jobs.go). Background work literally cannot preempt serving.
Why queries can't lock the AI DB
Rollups live in a separate ClickHouse store (omazy_analytics) with its own pool. Dashboards read pre-aggregated materialized views — millisecond point-reads, never a scan of the AI's Postgres.
Section 03 · the write path
Emit is cheap, and it swallows its own errors.
A producer fires a fact and moves on. Meter.Emit (internal/metering/meter.go) does a Redis INCRBY against a 60-second cached period snapshot (it never touches Postgres) plus an in-memory append — and it "never returns an error to the caller."
The ClickHouse insert is deferred to a 5-second flush loop, batched (200 rows for usage). If ClickHouse is slow or down, the batch spools to a durable Postgres analytics_outbox and a worker drains it every 30 seconds with exponential backoff — so a fact is never lost and a reply is never blocked. The rule for every producer: fire-and-forget, swallow-on-error, never await a flush on a request handler.
Section 04 · counter state & durability
Is the Redis counter persistent? What happens when Redis flushes?
Short answer: the Redis counter is a hot cache, not the ledger. A flush loses freshness, never data — and the counter self-heals.
flowchart LR
E["Meter.Emit"] -->|"leg 1 · fast"| R[("Redis counter
hot cache · TTL · day-keyed")]
E -->|"leg 2 · durable"| CH[("usage_events
ClickHouse · source of truth")]
CH --> RC["reconcile job
(nightly)"]
RC -->|"overwrite"| R
RC -->|"overwrite"| PG[("workspace_usage_counters
Postgres · reconciled totals")]
R -.->|"FLUSH · evict · TTL expiry"| X["counter lost"]
X -.->|"rebuilt on next reconcile"| R
usage_events row in ClickHouse (the source of truth). A reconcile job rebuilds Redis + the Postgres period totals from ClickHouse.Every increment lands in both Redis (fast, TTL'd, day-partitioned key) and the durable pipeline (usage_events in ClickHouse, with the analytics_outbox fallback). The two legs are independent, and a nightly reconcile job recomputes period totals from ClickHouse and overwrites Redis and the Postgres workspace_usage_counters.
| On a Redis flush / evict / TTL expiry… | Effect |
|---|---|
| Ledger / data integrity | None. Every increment is durably in usage_events regardless of Redis. Nothing is lost. |
| Fast-path limit gate (real-time quota) | Under-counts until the next reconcile → briefly fails open. Hard billing enforcement must read the reconciled total, not raw Redis. |
| Cheap dashboard tile reading Redis | Dips until reconcile. The authoritative number is always the ClickHouse rollup; treat Redis as a best-effort preview. |
Dedup keys (SET NX) | A returning visitor may count as new until the window passes. Bounded, non-financial; reconcile can de-dup from usage_events. |
The principle: never store anything in Redis you can't rebuild from ClickHouse. Redis persistence (RDB/AOF) is defense-in-depth, not a correctness dependency — we assume Redis can vanish and the reconcile restores it. This is also why a financial metric (§6) is forbidden from reading the raw counter.
Section 05 · cost-tiered access
You only pay for the stats you turn on.
Every metric declares a cost class — how expensive it is to produce. Cheap stats are free to everyone; live and streaming stats are opt-in and locked to higher plans. Real-time is never the default.
| Class | Produced by | Marginal cost | Default | Access |
|---|---|---|---|---|
| C0 cheap | Redis counter / cached snapshot | ~0 · O(1) | yes | all plans |
| C1 batch | materialized-view rollup | ~0 read | yes | all plans · daily→15-min by tier |
| C2 live | on-demand recompute + TTL cache, polled | 1 query / refresh | opt-in | Growth+ · 15–60s by tier |
| C3 stream | continuous recompute + fan-out + a subscriber per view | a running worker | opt-in | Business/Enterprise · per-widget |
No viewer, no compute
C3 subscriptions are ref-counted per (workspace, metric): the recompute loop + fan-out start on first subscribe and stop on last unsubscribe. An idle workspace runs zero dashboard background work.
A budget, not a cliff
C2/C3 draw from a per-workspace compute budget (metered like AI spend). Over budget → auto-downgrade to batch, never break. The resolver is the existing gate plus tierAllows(cost_class) ∧ withinComputeBudget.
Section 06 · the metric contract
A metric declares how it's made, how sensitive it is, and how fresh it must be.
Everything the pipeline needs to know about a stat is one registered object. Two fields do more work than they look: data_class (sensitivity, distinct from cost) and freshness (max_age_s + as_of, not a stored age).
type MetricDef = {
key: string // stable id, e.g. 'ai_conversations'
// -- production & cost (drives transport + gating) --
source: string // MV / counter / fact table it reads
rollup: 'sum'|'avg'|'count'|'p50'|'p95'|'distinct'|'ratio'
grain: 'counter'|'hour'|'day'
scope: 'workspace'|'app'
cost_class: 'C0'|'C1'|'C2'|'C3' // how expensive to produce
// -- classification (governance & source-of-truth) --
data_class: 'operational'|'financial'|'pii'|'sensitive'
// -- freshness (policy on the def; age is per-reading) --
max_age_s: number // acceptable staleness -> cache TTL + poll cadence
as_of: 'emit'|'rollup'|'reconcile' // which stage a value reflects
// -- access --
permission?: string // RBAC
min_tier?: PlanTier // entitlement (composes with cost_class)
realtime?: 'off'|'opt_in' // C3 default 'opt_in'
// -- refresh transport (derived from cost_class; NOT client-chosen) --
transport: 'snapshot'|'poll'|'stream'
refresh: Record<PlanTier, number> // seconds; per-tier poll cadence
} Do we keep data_class? Yes — it isn't cost_class.
cost_class answers how expensive; data_class answers what kind. They gate different things. A financial metric (revenue, refunds, ROI) or a pii-derived one needs stricter permissions and is audited on read; an operational one (sessions, response time) doesn't. Critically, data_class: 'financial' carries a hard rule — it must read the reconciled durable store, never the raw Redis counter (§4). One field, three jobs: access, audit, and source-of-truth.
Do we keep an age? Yes — as policy + provenance, not a stored number.
Don't put a mutable "age" on the definition. Model freshness two ways: max_age_s is a policy on the def (it drives the cache TTL and the poll cadence and when to show a "stale" marker), and as_of stamps each reading with the pipeline stage it reflects — so the UI can honestly say "as of 2m ago." Age is then just now − as_of, computed at read time, never persisted on the metric.
The refresh handler: HTTP, WebSocket, or Firebase?
The client never chooses — it declares the metric and range; the server decides transport from cost_class. So a widget's code doesn't change when a metric graduates from poll to stream.
| Cost class | Transport | Mechanism |
|---|---|---|
| C0 / C1 | HTTP snapshot | GET /metrics/snapshot once on load; cached. |
| C2 | HTTP poll | same endpoint, refetchInterval per tier (15–60s). No sockets. |
| C3 | WebSocket (web) · Ably (mobile) | the existing Redis-fanned realtime hub (/api/v1/realtime/connect), channel app:<id>:metrics, coalesced 1–2s, ref-counted. |
Firebase is not a metrics transport. FCM is push-notification delivery to the customer app — best-effort, throttled, device-token-addressed, unordered. Metric streaming needs a low-latency, ordered, in-session channel: that's the WebSocket hub on web and Ably on mobile, both of which the platform already runs. Using Firebase here would be the wrong tool.
Section 07 · web traffic & live chat
A dedicated lane for traffic; the realtime hub for live chat.
Web-traffic events get their own ingest path so they never ride the AI handler; the live-chat board rides the same realtime channels as the inbox — gated by cost class.
Web traffic — an isolated edge webhook
The widget beacons to a Cloudflare Worker at the edge, not an origin route — so when traffic hits the fan, the flood is absorbed, sampled, rate-limited and batched at the edge and never reaches the AI server's CPU. The origin pulls pre-batched events on its own schedule into the same outbox → ClickHouse spine. Traffic KPIs (visitors, sessions, MAU) are C0/C1 — free to every plan.
Live chat
"Who's chatting now / queue depth / agents online" is the C3 case on app:<id>:{inbox,metrics,presence} — Business/Enterprise, ref-counted. Growth tenants get the same board on a C2 poll. Its trends (volume, response time) are C1 rollups.
Section 08 · extensibility
Ship features by adding, never editing.
A new feature is an additive bundle: it emits new event types and registers new metric providers and widgets. The AI core, the event transport, and every existing dashboard stay untouched — so a launch can't break what's live.
flowchart LR
subgraph CORE["Core · untouched"]
ENV["versioned event envelope
(v · type · attrs Map)"]
REG["metric-provider + widget registries"]
AIP["AI serving path · existing dashboards"]
end
F["+ new feature: Payments"]
F -->|"emits"| EV["new event types
payment · refund"]
F -->|"registers"| MP["providers + MV
revenue · refund_rate"]
F -->|"registers"| WG["widgets
(each a cost_class)"]
EV --> ENV
MP --> REG
WG --> REG
Map attrs for open dimensions) means new dimensions need no ALTER; a CI diff-gate rejects any non-additive schema change.The launch superpower: because raw events are retained, a newly-registered metric is backfilled from history — you don't wait a month to collect data. The pattern mirrors the micro-app plugin platform already in the codebase (workspace_plugin_*, cai_ tokens): analytics providers are just another plugin capability.
Section 09 · implementation surface
Where it lives — and what's live vs net-new.
The write-path spine is on main. The dashboard/cost-tier layer is the design in the workspace plan.
Live today (middleware)
internal/metering/meter.go— non-blocking emit (Redis + outbox)internal/clickhouse/{buffered_writer,drain_worker,outbox}.go— batch + durable draininternal/jobs/jobs.go— Asynq queue isolation (chat off-queue)internal/billing/events.go— the durable bus pattern to generalizedb/clickhouse/migrations/*—usage_daily+*_5mrollups
Net-new (plan 11/12)
MetricDef+cost_class/data_class+ resolver predicate/ingest/eventscollector +web_eventsfact table- workspace
/metrics/{snapshot,timeseries}API app:<id>:metricsrealtime channel + ref-counted subscribe- wire the
messages/bot_responsesproducers (rollups have no rows yet)
Sources & further reading
The write-path spine is specified in RFC 2212; the dashboard engine, cost tiers, and extensibility are designed in the workspace platform plan:
docs/rfc2212-monetization-billing-metering.md— metering, counters, outbox, reconcile (the write path)docs/plan/workspace/12-telemetry-pipeline-and-extensibility.md— the pipeline, cost tiers, extensibility (this doc's full design)docs/plan/workspace/11-dashboards-and-widget-canvas.md— the widget canvas + saved views the metrics feedworkspace/docs/DASHBOARD_ENGINE.md— the in-repo engineering reference
Related engineering reading: Background Tasks (the Asynq platform this rides), the Common Channel Wrapper (where live-chat events originate), and the Agentic Harness.
// RFC 2212 + plan 11/12 · source in docs/
// edit → open a PR → ships on merge to main