Write a sentence, ship a workflow
An operator types: "When a VIP session opens, greet warmly and route to the senior agent." That sentence becomes a working automation in seconds. No graph to draw.
Omazy Engineering · Agentic Harness
A plain-English tour of the system that turns operator intent into customer outcomes — without graph builders, without code, without losing the audit trail.
Section 01 · history
Every era solved a real problem. Each one built the foundation for the next.
pattern matching
The first chatbot. Pretended to be a therapist by reflecting your sentence back. No understanding — just regex.
scripted bots
Commercial chatbots arrive. Hand-authored decision trees, ~10 million users. Cute, but brittle.
statistical NLU
First mainstream demonstration that machines could parse open-ended questions. Won Jeopardy.
intent + slots
The "intent classification" era. Operators map utterances → intents, intents → canned replies. Still scripted underneath.
RPA
Robotic Process Automation. Bots that click through enterprise software. Powerful, fragile, expensive.
visual workflows
The graph-builder boom. Drag boxes, connect arrows. Solved a real 2020 problem: operators can't code.
LLMs at scale
A statistical model that holds a conversation. Suddenly, chatbots could improvise.
tool-using LLMs
Claude Code, OpenAI Agents, Cline. The model picks tool calls. The harness becomes the platform.
skills, not graphs
Operators write a sentence in plain English. The harness loads it as a skill, the LLM picks tools, every step is traced.
Section 02 · the predecessor
A tour of the model the industry has used since 2020 — and what it gets right and wrong.
A typical n8n workflow. Every box is a step; every arrow is data; every JSON expression is a small piece of code.
n8n won the workflow market in 2020 because most operators couldn't write code. In 2026, most operators can describe what they want in English, and an LLM can pick the right tool. The graph builder solves a problem that no longer exists.
The DAG sketch isn't wrong — it's the wrong layer. Graph execution is the right model for deterministic data pipelines (ETL, finance reconciliation, compliance flows). It's the wrong model for CX, where every conversation is unique and the operator's intent is "respond well" not "execute these 7 specific steps."
We keep the DAG idea on the shelf for deterministic flows (the Phase B5 webhook chain, future ETL). We build this RFC for CX flows (chat, automations, business assistance). They share the same tool registry and trace store; they differ in how a "rule" is expressed and executed.
— RFC 2203, §1
Section 03 · why this matters
A harness that hides complexity from operators while preserving full control for engineers.
An operator types: "When a VIP session opens, greet warmly and route to the senior agent." That sentence becomes a working automation in seconds. No graph to draw.
The same engine answers customers in chat, runs automations in the background, and assists operators in their console. Skills move between surfaces.
The agent reads memory only through visible tool calls. Every decision is replayable. Nothing is silently injected into the prompt.
Postgres, Redis, OpenSearch — same operational surface as the rest of the platform. No new vendor to onboard, no new cost line on the invoice.
Section 04 · vocabulary
Twelve terms that show up everywhere. Worth a one-line definition each.
The runtime that runs the agent. One process, multi-tenant.
A markdown instruction file. The operator's unit of authoring.
An always-on guardrail. Brand voice, compliance, anti-hallucination.
Everything the LLM sees on a single call. Built fresh per run.
A capability the LLM can invoke. Built-in or vendor-supplied (MCP).
user_chat · autopilot · business_assist. The discriminator.
One end-to-end harness invocation. Tied to an audit trail.
The full record of a run. Append-only, replayable.
Durable knowledge across runs. Read/written through tools.
A lifecycle callback (pre_run, post_reply, etc.). Extension surface.
A child run inside a parent run. Used for task decomposition.
A symbolic permission like chats:write. Tools enforce it.
Section 05 · architecture
Three input surfaces converge through a single dispatch into the harness loop, and out to the platform's persistent stores.
↻ THE HARNESS LOOP
load context · call LLM · execute tool calls · feed back · until done
Section 06 · context
Eight layers, assembled fresh per run. The first four are cached on the model side; the last four vary every time.
The Omazy harness builds the context fresh every time it speaks to the model. Static parts — workspace identity, rules, mode guidance, the active skills — sit on top and get cached server-side, so the model doesn't pay for them again.
Variable parts — memory snippets the agent chose to read, the live conversation, the tools it's allowed to invoke right now, and the input that triggered this run — go at the bottom. The contract is simple: nothing enters the prompt unless it's accounted for in one of these eight layers.
Why this matters
Prompt caching cuts first-token latency by ~40% on warm runs. Operators get faster replies, the platform pays a smaller bill, and every layer is auditable from the trace.
Section 07 · scale
A single multi-tenant runtime serves every customer. Their differences live in data, not infrastructure.
SHARED RUNTIME
harness.Runner
stateless · multi-tenant · one Go process
Adding a new workspace is a database row, not a new server. Skills, rules, tools, memory, and budgets all live as per-tenant configuration loaded into the same runtime on demand.
Section 08 · governance
Guardrails layer top-down. Each tier appends to the one above it. A workspace can only override a platform rule with a logged compliance waiver.
Omazy-shipped — applies to every workspace. PII, profanity, factuality.
no-pii-leak.mdno-profanity.mdfactuality-vs-kb.md Workspace-authored — applies across all of one organisation's apps.
acme-brand-voice.md App-scoped — applies only when running inside one app/brand.
acme-cards-tone.md Per-mode tightening — chat mode might forbid what autopilot allows.
chat-no-billing-changes.md Rules append, never silently override. The compliance trail is queryable: every waiver is one row in rule_waivers.
Section 09a · memory
Each layer has a purpose, a store, and a TTL.
Section 09b · tools
Built-in tools ship with the harness. Vendor tools arrive through the marketplace.
BUILT-IN — every workspace
Chats
tag_sessionassign_toescalate_to_humansend_followupattach_form_to_replyComms
send_emailsend_smssend_pushKB
kb_searchkb_read_docCRM
customer_lookupcustomer_historyMemory
memory_readmemory_writememory_searchmemory_summarizeWorkflow
schedule_followupwait_for_eventsub_agent_callHuman
human_approval (suspend → operator click)VENDOR — installed via marketplace (RFC 1900)
Shopify
lookup_ordercreate_refundcheck_inventoryStripe
get_invoicerefund_chargelist_subscriptionsZendesk
create_ticketfind_similarset_statussend_templateopen_session_windowSlack
post_messagelookup_channelSection 10 · partner integration
A two-way sync layer that lets workspaces share skill libraries and MCP credentials with aicoach.pw — the operator-coaching platform.
SOURCE OF TRUTH FOR PLAYBOOKS
RUNTIME · MANY WORKSPACES
The adapter sits between the two systems and treats AICoach as a content authority for generic playbooks (e.g., "how to handle a refund request" for an e-commerce vertical). Omazy workspaces inherit those skills automatically, layer their own brand-specific overrides on top, and stream anonymised trace data back so AICoach can improve the next version.
// workspaces opt-in per skill bundle. Overrides always win locally. No telemetry without consent.
Section 11 · the new chat flow
A complete customer chat, end to end.
Section 12 · glossary of decisions
Each one is a non-obvious choice. If anyone disagrees, RFC 2203 is back open.
// sources: docs/rfc2203-agentic-harness.md
// related: technical deep-dive →