Omazy Engineering
How we build it.
Living engineering documentation. Architecture decisions, RFCs, and platform maps for the team — and anyone curious how Omazy is built.
13 surfaces · click to open
A model swap is a prompt change you did not write
One Prompt, Three Models
Two providers went dark on the same day and a whole agent fleet slid onto its last fallback with nothing failing to deploy. The findings that came out of getting back: why running out of money arrives as a 429 from one vendor and a 400 from another, what a fallback chain costs per request without a circuit breaker, the preflight that passed and still let through a model four times slower, the five places a reasoning model breaks a driver built for chat, and how three model families behaved against one strict system prompt — including the one that invented a document owner and tagged it with the highest confidence value in the vocabulary.
A page is not an inbox
The Channel With No Webhook
LinkedIn Pages as a first-class channel: why a comment becomes an ordinary inbound message, the optional driver interface for the things a page can do that an inbox cannot, and an OAuth flow made deliberately longer because a reviewer has to watch a user pick their own Page. Then the number that turned out to be the design: 500 requests per day shared across every tenant, which makes the polling interval a division rather than a setting. Ends with the bug that would have published a support reply to every follower.
The green results that lied
Building a Workflow Engine
A node-graph automation engine built on top of a durable run engine that already existed, because a rule and a task are both degenerate graphs. The return type that has to admit a third of node kinds never finish, checkpoint-before-dispatch, fencing tokens against at-least-once delivery, and the single-flight key that starves an event-triggered workflow. Then the part worth reading twice: four separate occasions where a test came back green while proving nothing, including an empty page passing every layout check and a regression test that passed with the fix reverted.
Size it before you build it
Data Export Engine
Getting a workspace out of the platform is a file-handling problem: a contact list might be forty rows or four hundred thousand, and the same button produces both. Estimating from an indexed count and a per-dataset row width, routing small single exports into the request while everything larger goes to a worker, and an hourly sweep that deletes the object before it flips the row. Measured on production catalogs from 146 to 14,881 rows, including where the estimate is wrong by 3.5x and why that is the safe direction.
Authored top down, resolved bottom up
Voice Configuration & Determinism
Reading voice configuration at process start freezes the authoring order at boot: one process, one tenant, one voice, and a table the console writes that nothing reads at call time. Per-call resolution with an environment fallback, a cache namespace hashed over the phrase text so stale audio is unreachable by construction, keyword routing that answers in 0.75 s against 3.50 s for a generated turn, and why changing a voice is a priced operation rather than a dropdown.
Latency, measured on live calls
Making a Voice Agent Fast
From a phone agent that connected and played silence to a greeting that starts in 820 ms. Four separate causes of a silent call, a turn loop with no memory, a nil pointer a full green test suite could not see, detector retuning for Bengali, and the measured budget where speech synthesis is 59 percent of the wait and costs 12.2 ms per character. Includes the two experiments that made it slower.
Platform map
Know Omazy
Every workspace, app, agent, and admin module — surfaces, statuses, and the launch roadmap. Filter by status, scope, search by name.
How the agentic system works
Omazy Harness
A plain-English tour of the runtime that turns operator intent into customer outcomes. Architecture, vocabulary, memory, tools, and the new chat flow — visualised.
Marketplace + extensibility
Omazy Plugins
How vendors extend Omazy with MCP-compatible plugins, hosted on their own infra, with Stripe Connect billing baked in. Five capability layers, twelve trust controls.
Automations + agentic harness
Agentic Harness — deep dive
The technical companion. Three flow types — simple chat, deterministic automation (RFC 2202), probabilistic agentic harness (RFC 2203). Skills, not graphs.
Asynq + Redis platform
Background Tasks
How every email, export, automation, and analytics drain rides one queue platform. Architecture, lifecycle, monitoring, gaps — and a load matrix vs Cloudflare Queues and RabbitMQ.
Omnichannel engine + CIR
Common Channel Wrapper
How WhatsApp, Telegram, Instagram, Email and a dozen more collapse into one canonical pipeline: a stateless driver engine, a Common International Representation for messages, per-config webhook rotation, and a business-owner Channel Manager.
Metrics without slowing the AI core
Dashboard & Telemetry Engine
A three-plane pipeline — emit → process → serve — where the cost of producing a stat decides who sees it (real-time is opt-in, tier-locked), counters survive a Redis flush, and shipping a feature can never break what is live.
// principles
How we make calls
The non-obvious rules that shape every architecture review and PR comment. Older than any one RFC; newer ones get written down here first.
- ·01
Driver pattern for every external vendor
Email, SMS, payments, AI providers — each behind a narrow interface + factory. Swapping a vendor is one impl file + one factory case.
- ·02
No new infra primitives without a customer
Go + Asynq + Postgres + Redis + R2 + OpenSearch + ClickHouse. We don't add Temporal, LangChain, or Kubernetes until a real workload breaks the existing stack.
- ·03
One process, per-tenant config
Workspaces are rows in PG and blobs in R2. No per-business container. Isolation is enforced at the data boundary.
- ·04
RFC-driven decisions
Big architectural moves get an RFC + open issues + an explicit roadmap. Small fixes ship; load-bearing decisions get reviewed in writing.
- ·05
Instructions, not graphs
Operators describe intent in markdown. The LLM picks tool calls. Graph builders are reserved for deterministic ETL — wrong layer for CX.
- ·06
Trace everything, replay anything
Every agent run is append-only and full-fidelity. PG for hot reads, ClickHouse for analytics. Foundation for offline eval and fine-tuning.
// sources live in docs/ on each repo
// edit a page → open a PR → it ships on merge to main