Skip to main content
Omazy Engineering

RFC 2212 · Telemetry & dashboards

Measure everything. Slow nothing.

How the workspace dashboards collect, store and serve every metric behind a hard boundary from the AI core — where the cost of producing a stat decides who can see it, and shipping a feature can never break what is live.

// three planes · cost tiers · additive-only // status: metering spine live · dashboard engine in design (plan 11/12)

Section 01 · the problem

A dashboard that lags reality is useless. One that slows the AI is worse.

Business owners want to watch their business live — conversations, revenue, ROI. But the AI serving path is sacred: a chat reply must never wait on a metrics write, and a heavy dashboard query must never lock a table the agent loop needs.

So the whole system is built around one boundary. The AI core does exactly one measurement thing — a sub-millisecond counter bump — and everything else happens later, on separate workers, against a separate database. On top of that we layer a second rule the industry usually forgets: producing a stat costs money, so what it costs decides who can see it. Real-time is opt-in and tier-locked, never a flat tax on every request.

Section 02 · architecture

Three planes, one hard boundary.

Emit is the only thing on the request path, and it is O(1). Process and Serve run entirely off it — a metrics backlog or a ClickHouse outage is invisible to a customer mid-chat.

flowchart LR
  subgraph HOT["① EMIT · request path · O(1)"]
    AI["AI core / chat reply"] -->|"Redis INCR (cached snapshot)"| R[("Redis counters")]
    AI -->|"append"| B["in-proc buffer"]
    WDG["Widget / website"] -->|"sendBeacon"| COL["edge collector · Worker
(separate from origin CPU)"] end subgraph PROC["② PROCESS · isolated workers (Asynq low)"] B -->|"batch 200/5s"| W["buffered writers"] COL --> W W -->|"on CH failure"| OB[("analytics_outbox · Postgres")] OB -->|"drain @30s · SKIP LOCKED"| CH[("ClickHouse · omazy_analytics")] W --> CH CH --> MV[("materialized views · rollups")] end subgraph SERVE["③ SERVE · read-only"] MV --> API["/metrics API · ms point-reads"] R --> API API --> D["dashboard widgets"] MV -.->|"opt-in · ref-counted"| PUB["realtime publisher (C3)"] PUB -.-> D end
Fig 1 — ① Emit (request path) → ② Process (isolated workers) → ③ Serve (read-only rollups). The realtime publisher is opt-in and ref-counted, not always-on.

Why chat can't be starved

Chat/AI is synchronous HTTP/SSE — it never enters the job queue. All analytics runs on Asynq's weighted low queue (internal/jobs/jobs.go). Background work literally cannot preempt serving.

Why queries can't lock the AI DB

Rollups live in a separate ClickHouse store (omazy_analytics) with its own pool. Dashboards read pre-aggregated materialized views — millisecond point-reads, never a scan of the AI's Postgres.

Section 03 · the write path

Emit is cheap, and it swallows its own errors.

A producer fires a fact and moves on. Meter.Emit (internal/metering/meter.go) does a Redis INCRBY against a 60-second cached period snapshot (it never touches Postgres) plus an in-memory append — and it "never returns an error to the caller."

The ClickHouse insert is deferred to a 5-second flush loop, batched (200 rows for usage). If ClickHouse is slow or down, the batch spools to a durable Postgres analytics_outbox and a worker drains it every 30 seconds with exponential backoff — so a fact is never lost and a reply is never blocked. The rule for every producer: fire-and-forget, swallow-on-error, never await a flush on a request handler.

Section 04 · counter state & durability

Is the Redis counter persistent? What happens when Redis flushes?

Short answer: the Redis counter is a hot cache, not the ledger. A flush loses freshness, never data — and the counter self-heals.

flowchart LR
  E["Meter.Emit"] -->|"leg 1 · fast"| R[("Redis counter
hot cache · TTL · day-keyed")] E -->|"leg 2 · durable"| CH[("usage_events
ClickHouse · source of truth")] CH --> RC["reconcile job
(nightly)"] RC -->|"overwrite"| R RC -->|"overwrite"| PG[("workspace_usage_counters
Postgres · reconciled totals")] R -.->|"FLUSH · evict · TTL expiry"| X["counter lost"] X -.->|"rebuilt on next reconcile"| R
Fig 2 — Emit writes two legs: a fast Redis counter and a durable usage_events row in ClickHouse (the source of truth). A reconcile job rebuilds Redis + the Postgres period totals from ClickHouse.

Every increment lands in both Redis (fast, TTL'd, day-partitioned key) and the durable pipeline (usage_events in ClickHouse, with the analytics_outbox fallback). The two legs are independent, and a nightly reconcile job recomputes period totals from ClickHouse and overwrites Redis and the Postgres workspace_usage_counters.

On a Redis flush / evict / TTL expiry…Effect
Ledger / data integrityNone. Every increment is durably in usage_events regardless of Redis. Nothing is lost.
Fast-path limit gate (real-time quota)Under-counts until the next reconcile → briefly fails open. Hard billing enforcement must read the reconciled total, not raw Redis.
Cheap dashboard tile reading RedisDips until reconcile. The authoritative number is always the ClickHouse rollup; treat Redis as a best-effort preview.
Dedup keys (SET NX)A returning visitor may count as new until the window passes. Bounded, non-financial; reconcile can de-dup from usage_events.

The principle: never store anything in Redis you can't rebuild from ClickHouse. Redis persistence (RDB/AOF) is defense-in-depth, not a correctness dependency — we assume Redis can vanish and the reconcile restores it. This is also why a financial metric (§6) is forbidden from reading the raw counter.

Section 05 · cost-tiered access

You only pay for the stats you turn on.

Every metric declares a cost class — how expensive it is to produce. Cheap stats are free to everyone; live and streaming stats are opt-in and locked to higher plans. Real-time is never the default.

ClassProduced byMarginal costDefaultAccess
C0 cheapRedis counter / cached snapshot~0 · O(1)yesall plans
C1 batchmaterialized-view rollup~0 readyesall plans · daily→15-min by tier
C2 liveon-demand recompute + TTL cache, polled1 query / refreshopt-inGrowth+ · 15–60s by tier
C3 streamcontinuous recompute + fan-out + a subscriber per viewa running workeropt-inBusiness/Enterprise · per-widget

No viewer, no compute

C3 subscriptions are ref-counted per (workspace, metric): the recompute loop + fan-out start on first subscribe and stop on last unsubscribe. An idle workspace runs zero dashboard background work.

A budget, not a cliff

C2/C3 draw from a per-workspace compute budget (metered like AI spend). Over budget → auto-downgrade to batch, never break. The resolver is the existing gate plus tierAllows(cost_class) ∧ withinComputeBudget.

Section 06 · the metric contract

A metric declares how it's made, how sensitive it is, and how fresh it must be.

Everything the pipeline needs to know about a stat is one registered object. Two fields do more work than they look: data_class (sensitivity, distinct from cost) and freshness (max_age_s + as_of, not a stored age).

type MetricDef = {
  key: string                          // stable id, e.g. 'ai_conversations'

  // -- production & cost (drives transport + gating) --
  source: string                       // MV / counter / fact table it reads
  rollup: 'sum'|'avg'|'count'|'p50'|'p95'|'distinct'|'ratio'
  grain:  'counter'|'hour'|'day'
  scope:  'workspace'|'app'
  cost_class: 'C0'|'C1'|'C2'|'C3'      // how expensive to produce

  // -- classification (governance & source-of-truth) --
  data_class: 'operational'|'financial'|'pii'|'sensitive'

  // -- freshness (policy on the def; age is per-reading) --
  max_age_s: number                    // acceptable staleness -> cache TTL + poll cadence
  as_of:     'emit'|'rollup'|'reconcile'   // which stage a value reflects

  // -- access --
  permission?: string                  // RBAC
  min_tier?:   PlanTier                // entitlement (composes with cost_class)
  realtime?:   'off'|'opt_in'          // C3 default 'opt_in'

  // -- refresh transport (derived from cost_class; NOT client-chosen) --
  transport: 'snapshot'|'poll'|'stream'
  refresh:   Record<PlanTier, number>  // seconds; per-tier poll cadence
}

Do we keep data_class? Yes — it isn't cost_class.

cost_class answers how expensive; data_class answers what kind. They gate different things. A financial metric (revenue, refunds, ROI) or a pii-derived one needs stricter permissions and is audited on read; an operational one (sessions, response time) doesn't. Critically, data_class: 'financial' carries a hard rule — it must read the reconciled durable store, never the raw Redis counter (§4). One field, three jobs: access, audit, and source-of-truth.

Do we keep an age? Yes — as policy + provenance, not a stored number.

Don't put a mutable "age" on the definition. Model freshness two ways: max_age_s is a policy on the def (it drives the cache TTL and the poll cadence and when to show a "stale" marker), and as_of stamps each reading with the pipeline stage it reflects — so the UI can honestly say "as of 2m ago." Age is then just now − as_of, computed at read time, never persisted on the metric.

The refresh handler: HTTP, WebSocket, or Firebase?

The client never chooses — it declares the metric and range; the server decides transport from cost_class. So a widget's code doesn't change when a metric graduates from poll to stream.

Cost classTransportMechanism
C0 / C1HTTP snapshotGET /metrics/snapshot once on load; cached.
C2HTTP pollsame endpoint, refetchInterval per tier (15–60s). No sockets.
C3WebSocket (web) · Ably (mobile)the existing Redis-fanned realtime hub (/api/v1/realtime/connect), channel app:<id>:metrics, coalesced 1–2s, ref-counted.

Firebase is not a metrics transport. FCM is push-notification delivery to the customer app — best-effort, throttled, device-token-addressed, unordered. Metric streaming needs a low-latency, ordered, in-session channel: that's the WebSocket hub on web and Ably on mobile, both of which the platform already runs. Using Firebase here would be the wrong tool.

Section 07 · web traffic & live chat

A dedicated lane for traffic; the realtime hub for live chat.

Web-traffic events get their own ingest path so they never ride the AI handler; the live-chat board rides the same realtime channels as the inbox — gated by cost class.

Web traffic — an isolated edge webhook

The widget beacons to a Cloudflare Worker at the edge, not an origin route — so when traffic hits the fan, the flood is absorbed, sampled, rate-limited and batched at the edge and never reaches the AI server's CPU. The origin pulls pre-batched events on its own schedule into the same outbox → ClickHouse spine. Traffic KPIs (visitors, sessions, MAU) are C0/C1 — free to every plan.

Live chat

"Who's chatting now / queue depth / agents online" is the C3 case on app:<id>:{inbox,metrics,presence} — Business/Enterprise, ref-counted. Growth tenants get the same board on a C2 poll. Its trends (volume, response time) are C1 rollups.

Section 08 · extensibility

Ship features by adding, never editing.

A new feature is an additive bundle: it emits new event types and registers new metric providers and widgets. The AI core, the event transport, and every existing dashboard stay untouched — so a launch can't break what's live.

flowchart LR
  subgraph CORE["Core · untouched"]
    ENV["versioned event envelope
(v · type · attrs Map)"] REG["metric-provider + widget registries"] AIP["AI serving path · existing dashboards"] end F["+ new feature: Payments"] F -->|"emits"| EV["new event types
payment · refund"] F -->|"registers"| MP["providers + MV
revenue · refund_rate"] F -->|"registers"| WG["widgets
(each a cost_class)"] EV --> ENV MP --> REG WG --> REG
Fig 3 — A feature plugs into three registries. The versioned, additive-only envelope (Map attrs for open dimensions) means new dimensions need no ALTER; a CI diff-gate rejects any non-additive schema change.

The launch superpower: because raw events are retained, a newly-registered metric is backfilled from history — you don't wait a month to collect data. The pattern mirrors the micro-app plugin platform already in the codebase (workspace_plugin_*, cai_ tokens): analytics providers are just another plugin capability.

Section 09 · implementation surface

Where it lives — and what's live vs net-new.

The write-path spine is on main. The dashboard/cost-tier layer is the design in the workspace plan.

Live today (middleware)

  • internal/metering/meter.go — non-blocking emit (Redis + outbox)
  • internal/clickhouse/{buffered_writer,drain_worker,outbox}.go — batch + durable drain
  • internal/jobs/jobs.go — Asynq queue isolation (chat off-queue)
  • internal/billing/events.go — the durable bus pattern to generalize
  • db/clickhouse/migrations/* — usage_daily + *_5m rollups

Net-new (plan 11/12)

  • MetricDef + cost_class/data_class + resolver predicate
  • /ingest/events collector + web_events fact table
  • workspace /metrics/{snapshot,timeseries} API
  • app:<id>:metrics realtime channel + ref-counted subscribe
  • wire the messages/bot_responses producers (rollups have no rows yet)

Sources & further reading

The write-path spine is specified in RFC 2212; the dashboard engine, cost tiers, and extensibility are designed in the workspace platform plan:

  • docs/rfc2212-monetization-billing-metering.md — metering, counters, outbox, reconcile (the write path)
  • docs/plan/workspace/12-telemetry-pipeline-and-extensibility.md — the pipeline, cost tiers, extensibility (this doc's full design)
  • docs/plan/workspace/11-dashboards-and-widget-canvas.md — the widget canvas + saved views the metrics feed
  • workspace/docs/DASHBOARD_ENGINE.md — the in-repo engineering reference

Related engineering reading: Background Tasks (the Asynq platform this rides), the Common Channel Wrapper (where live-chat events originate), and the Agentic Harness.

// RFC 2212 + plan 11/12 · source in docs/

// edit → open a PR → ships on merge to main