Skip to main content
Omazy Engineering Engineering · RFC 2202 + RFC 2203

Agentic Harness & Automations

How Omazy runs CX work — three flow types, one harness, one tool registry. Skills, not graphs.

RFC 2202 + RFC 2203 Draft Updated 2026-05-08

Three flow types, one platform

Different shapes of work need different execution models. We ship all three. Pick by complexity + cost — not by team preference.

ONE LLM CALL

Simple chat

Customer asks → bot answers

A single LLM invocation against grounded context (persona + KB excerpt + recent thread). No tool loops, no branching, no memory writes.

WHERE

Default for /chat replies that don't need external lookups.

$ — cheapest, ~500ms median, sub-3s SLA

RFC 2202 CHAIN

Deterministic automation

Event → match → run actions in order

Imperative actions[] chain. Strict ordering, no LLM reasoning, no branching. Each action is a Go handler.

WHERE

When you know exactly what should happen — strict ETL, compliance flows, fixed notifications.

¢ — near-zero compute, no LLM tokens

RFC 2203 HARNESS

Probabilistic automation

Skill describes intent → LLM picks tool calls

Markdown skill file + frontmatter. Harness loop builds context, calls LLM, executes tool calls, repeats until the LLM emits a final reply.

WHERE

CX work that benefits from reasoning, memory, branching, tone — chat, autopilot triage, operator copilot.

$$ — LLM cost per call, ~2–15s, budget-capped

Customer message arrives
via web widget / WhatsApp / email
chats domain stores message
session.messages append
Build context (cheap)
agent persona + KB excerpt + last 10 turns
Single LLM call
Anthropic / OpenAI via driver-pattern gateway
Stream tokens to customer
SSE through chats reply pipeline

Definition · grounded context

Persona (agent system prompt) + knowledge-base excerpt (top-k retrieval against bot_documents) + recent thread. Built fresh per call. No agent memory, no tool calls.

Why no tool loop here

90% of replies are pure Q&A. A tool loop adds latency (2nd round-trip ≈ +1.5s) and cost without changing the answer. When intent NEEDS a tool (lookup_order, refund), chats escalates the same request to the harness in user_chat mode — same code path, switched on intent classification.

Engineering decision · cache the system prompt

Persona + KB excerpt is mostly static across a session. Anthropic prompt-caching keeps it warm — first-token latency drops ~40% on cache hits. Cache-invalidate on KB reindex.

SLA budget

First token < 1s · full reply < 3s · 99p < 8s (measured per-app; budget alarms at p95 breach).

Producer event
chats / agents / KB / members / budgets
Publisher.Publish(workspace, event, payload)
non-fatal — never blocks producer
Match active rules · apply filters
WHERE event=? AND active AND filters@>
Enqueue Asynq · TaskAutomationExecute
queue=default · retry=3 · retention=24h
↻ LOOP
Runner walks actions[] in order
send_email → assign_to → tag_session
Persist run + step_results to PG
automation_runs row · step_results JSONB

Definition · deterministic

Same event + same rule = same action chain executed in the same order. No reasoning. Failure modes are predictable (network errors, idempotency violations) — not "wrong tool".

Why a flat array (not a graph)

95% of automation rules in CRM/CX are 1–4 steps in linear order: "send email, assign, tag, escalate." A graph builder for that is yak shaving. The DAG sketch (RFC 2202 §15) is kept on the shelf for ETL-style flows where branching/data-passing matter — not for CX automations, where the LLM harness (Flow C) is a better fit when complexity grows.

Engineering decision · server-side filtering

Filters (app_handles, tags) eval in Publisher.Publish BEFORE Asynq enqueue. A workspace with 50 rules where 2 match → 2 tasks queued, not 50. Queue depth is a real cost; filter eval is one indexed JSONB query.

Engineering decision · skipped ≠ failed

Stub action handlers return StepStatusSkipped (not failed). Run aggregates as "succeeded" if all real steps succeed, letting operators ship full chains today.

user_chat mode
autopilot mode
business_assist mode
harness.Dispatch
resolve mode + actor + capabilities
Build context pack
skills + rules + memory + tools + input
pre_run hooks
budget · PII redact · rule load

↻ THE HARNESS LOOP

internal/harness/runner.go

load ctx → pre_run → LLM call → tool calls →

post_tool → ( back to LLM call ) → post_reply → done

max_steps=10 · budget enforced · every step traced

post_reply hooks
factuality · brand voice · profanity
Persist trace + memory writes
agent_runs + agent_run_steps · customer_memory

Definition · agentic harness

A loop that calls an LLM with tool definitions, executes the tools the LLM picks, and feeds results back until the LLM emits a final answer. Same shape as Claude Code, OpenAI Agents SDK, Cline.

Why one harness, three modes

user_chat / autopilot / business_assist all share the SAME loop. They differ in which skills are loaded, which tools are exposed, which guardrails apply, whose identity is on the wire. Mode is config — not a separate codebase. Cuts duplicated tool-handler/trace/budget plumbing by 3×.

Why instructions, not graphs

Skill = markdown + YAML frontmatter ("when VIP session opens, greet warmly + route + ping #vip"). LLM reads it, picks tool order. n8n/Flowise solve a 2020 problem (operators couldn't code) — in 2026 operators describe intent in English and an LLM picks tools. Graphs are for ETL, not CX.

Engineering decision · explicit memory R/W

Memory is read/written through tools (memory_read, memory_write), never silently injected into the prompt. Auditability: trace shows exactly which memories shaped the output. Bounded context: LLM picks what it needs, not an opinionated retriever stuffing 8K tokens.

Engineering decision · capability boundary

Tools — not the LLM — enforce permissions. Even if the LLM "tries" to call a forbidden tool, the registry rejects and returns a structured error. Capabilities are the security perimeter, not prompt instructions.

Engineering decision · max_steps budget

Default 10 LLM-calls per run. Pre-tool hook rejects redundant calls (same args twice). Trace surfaces step explosion. Stops runaway loops cold.

Skills

instruction store

  • markdown + frontmatter
  • stored in PG + R2
  • selected per mode

Rules

always-on guardrails

  • platform → workspace →
  • app → mode (layered)
  • precedence-ordered

Memory

explicit R/W

  • conversation (Redis)
  • customer (PG + pgvector)
  • workspace (PG)

Tools

callable capabilities

  • built-in (4 from RFC 2202)
  • + vendor MCP (RFC 1900)
  • per-workspace allowlist

CONTEXT PACK

typed Go struct · built fresh per run

↻ THE HARNESS LOOP

internal/harness/runner.go

load ctx → pre_run → LLM call → tool calls →

post_tool → ( back to LLM call ) → post_reply → done

max_steps=10 · budget enforced · every step traced

LLM Gateway

driver pattern

  • Anthropic primary
  • OpenAI fallback
  • budget enforcement

Hooks

extension surface

  • pre_run / pre_tool /
  • post_tool / pre_reply /
  • post_reply (Claude Code)

Trace store

append-only

  • agent_runs (PG)
  • agent_run_steps (PG+CH)
  • full replay-ability

Capabilities

security perimeter

  • chats:write · mail:send
  • vendor:shopify:read …
  • enforced at tool boundary

CUSTOMER-FACING

user_chat

TRIGGER
inbound chat message
ACTS AS
workspace (system) acting on behalf of agent
TOOLS
read-mostly + safe writes
(NO billing changes)
GUARDRAILS
STRICT — anti-hallucination,
brand tone, profanity, PII
MEMORY
conversation + customer profile
SLA
first token <1s · reply <3s
REPLACES
the ad-hoc chats reply pipeline

BUSINESS BACKGROUND

autopilot

TRIGGER
RFC 2202 Publisher event
or scheduled cron
ACTS AS
workspace (system)
TOOLS
full configured tool set
+ vendor MCP
GUARDRAILS
MEDIUM — idempotency,
budget caps, vendor rate-limits
MEMORY
workspace + run history
("don't email twice today")
SLA
no streaming · run <5min
REPLACES
RFC 2202 v1 chains (opt-in v3)

OPERATOR COPILOT

business_assist

TRIGGER
operator UI action
(summarize / draft / triage)
ACTS AS
operator (their identity)
TOOLS
full operator set
+ admin tools
GUARDRAILS
LIGHT — operator reviews
before applying
MEMORY
session scratchpad +
operator recent activity
SLA
first token <1s · streaming
REPLACES
(new — no equivalent today)
01

Instructions, not graphs

  • Skill = markdown + YAML frontmatter
  • LLM picks tool order at runtime
  • DAG sketch (RFC 2202 §15) reserved for ETL, not CX
02

One harness, three modes

  • user_chat / autopilot / business_assist share the same loop
  • Modes differ only in skills, tools, guardrails, identity
  • Adding a mode = rows + a tool allowlist (not a new codebase)
03

One process, per-tenant config

  • No per-business container
  • Workspaces are rows in PG + R2 blobs
  • Scaling = scaling the harness process
  • Isolation enforced at the data boundary
04

Memory is explicit, not auto-injected

  • memory_read / memory_write / memory_search tools
  • Nothing silently appended to the prompt
  • Trace shows exactly which memories shaped the output
05

Tools = security boundary

  • Capabilities enforced at the Tool.Execute boundary
  • Not enforced by prompt instructions
  • LLM trying to call a forbidden tool gets a structured error
06

RFC 2202 stays — v3 adds skill_id

  • v1 chains keep running unchanged
  • New automations default to v3 (skill_id reference)
  • "Promote to skill" is opt-in, no forced migration
  • Same automation_runs row schema
07

Driver pattern for LLM provider

  • Anthropic primary, OpenAI fallback
  • Same pattern as email / SMS drivers
  • Caller never branches on provider
  • Swap = one factory case + one impl file
08

Prompt caching as a first-class concern

  • Skill body + rule body + persona cached server-side
  • Per-run variation (memory + input) goes last
  • Cuts first-token latency by ~40% on cache hits
09

Trace is append-only, full-fidelity

  • PG for hot reads, ClickHouse mirror for analytics
  • Every run replayable from stored inputs
  • Foundation for offline eval + future fine-tuning
10

Hooks borrowed from Claude Code

  • pre_run / pre_tool / post_tool / pre_reply / post_reply
  • Workspaces register Go funcs or sandboxed shell hooks
  • Familiar shape for engineers + operators
11

AI authors skills, not just runs them

  • business_assist mode includes an "author-skill" skill
  • Operators describe a rule in English
  • Copilot drafts the skill.md; operator reviews + saves
  • Scales rule authoring without a 50-step wizard
12

No new infra primitives

  • Go + Asynq + Postgres + Redis + R2
  • OpenSearch (search + vector) · ClickHouse (analytics)
  • Existing RFC 1900 MCP gateway
  • No Temporal, LangChain, or graph engine
  • The loop is ~200 lines of Go we own

SOURCES

  • docs/rfc2202-automations.md ·
  • docs/rfc2203-agentic-harness.md ·
  • docs/rfc2201-platform-core.md ·
  • docs/design/agentic-infographic.pen

Engine: Go + Anthropic SDK · Asynq · Postgres + pgvector · OpenSearch (search + vector) · ClickHouse (trace analytics) · Redis · R2 · MCP gateway (RFC 1900). No new infra primitives.