Skip to main content
Omazy Engineering

Omazy Engineering · Agentic Harness

How Omazy runs the work behind every chat.

A plain-English tour of the system that turns operator intent into customer outcomes — without graph builders, without code, without losing the audit trail.

Section 01 · history

A short history of chatbots and automation.

Every era solved a real problem. Each one built the foundation for the next.

  1. 1966

    pattern matching

    ELIZA

    The first chatbot. Pretended to be a therapist by reflecting your sentence back. No understanding — just regex.

  2. 2001

    scripted bots

    SmarterChild (AIM)

    Commercial chatbots arrive. Hand-authored decision trees, ~10 million users. Cute, but brittle.

  3. 2011

    statistical NLU

    IBM Watson

    First mainstream demonstration that machines could parse open-ended questions. Won Jeopardy.

  4. 2014

    intent + slots

    Dialogflow / Wit.ai

    The "intent classification" era. Operators map utterances → intents, intents → canned replies. Still scripted underneath.

  5. 2018

    RPA

    UiPath, Automation Anywhere

    Robotic Process Automation. Bots that click through enterprise software. Powerful, fragile, expensive.

  6. 2020

    visual workflows

    n8n · Zapier · Make

    The graph-builder boom. Drag boxes, connect arrows. Solved a real 2020 problem: operators can't code.

  7. 2022

    LLMs at scale

    ChatGPT

    A statistical model that holds a conversation. Suddenly, chatbots could improvise.

  8. 2024

    tool-using LLMs

    Agent SDKs

    Claude Code, OpenAI Agents, Cline. The model picks tool calls. The harness becomes the platform.

  9. 2026

    skills, not graphs

    Omazy Agentic Harness

    Operators write a sentence in plain English. The harness loads it as a skill, the LLM picks tools, every step is traced.

Section 02 · the predecessor

How n8n (and the workflow-builder family) works.

A tour of the model the industry has used since 2020 — and what it gets right and wrong.

trigger Webhook /orders/created
action HTTP Request GET /customer/{{$json.id}}
if VIP? {{ $json.tier === 'vip' }}
true Slack #vip-alerts
false Set tag = "standard"

A typical n8n workflow. Every box is a step; every arrow is data; every JSON expression is a small piece of code.

What graph builders solved.

  • + Operators stopped waiting on engineers for simple integrations.
  • + Visible flow: anyone could read the graph and understand the logic.
  • + Errors were inspectable. Each node had inputs and outputs.

Where they hit the ceiling.

  • − Anything beyond a flowchart needed JavaScript inside boxes.
  • − Customer support is not a flowchart. Every conversation is unique.
  • − Adding a new branch meant dragging more nodes — exponential complexity.
EVOLUTION 2026
  1. TIMING

    n8n won the workflow market in 2020 because most operators couldn't write code. In 2026, most operators can describe what they want in English, and an LLM can pick the right tool. The graph builder solves a problem that no longer exists.

  2. LAYER

    The DAG sketch isn't wrong — it's the wrong layer. Graph execution is the right model for deterministic data pipelines (ETL, finance reconciliation, compliance flows). It's the wrong model for CX, where every conversation is unique and the operator's intent is "respond well" not "execute these 7 specific steps."

  3. BOTH

    We keep the DAG idea on the shelf for deterministic flows (the Phase B5 webhook chain, future ETL). We build this RFC for CX flows (chat, automations, business assistance). They share the same tool registry and trace store; they differ in how a "rule" is expressed and executed.

— RFC 2203, §1

Section 03 · why this matters

The value of removing complexity for the operator.

A harness that hides complexity from operators while preserving full control for engineers.

01

Write a sentence, ship a workflow

An operator types: "When a VIP session opens, greet warmly and route to the senior agent." That sentence becomes a working automation in seconds. No graph to draw.

02

One platform, three audiences

The same engine answers customers in chat, runs automations in the background, and assists operators in their console. Skills move between surfaces.

03

Memory the operator can audit

The agent reads memory only through visible tool calls. Every decision is replayable. Nothing is silently injected into the prompt.

04

Built on what we already trust

Postgres, Redis, OpenSearch — same operational surface as the rest of the platform. No new vendor to onboard, no new cost line on the invoice.

Section 04 · vocabulary

A short legend.

Twelve terms that show up everywhere. Worth a one-line definition each.

Harness

The runtime that runs the agent. One process, multi-tenant.

Skill

A markdown instruction file. The operator's unit of authoring.

Rule

An always-on guardrail. Brand voice, compliance, anti-hallucination.

Context Pack

Everything the LLM sees on a single call. Built fresh per run.

Tool

A capability the LLM can invoke. Built-in or vendor-supplied (MCP).

Mode

user_chat · autopilot · business_assist. The discriminator.

Run

One end-to-end harness invocation. Tied to an audit trail.

Trace

The full record of a run. Append-only, replayable.

Memory

Durable knowledge across runs. Read/written through tools.

Hook

A lifecycle callback (pre_run, post_reply, etc.). Extension surface.

Sub-agent

A child run inside a parent run. Used for task decomposition.

Capability

A symbolic permission like chats:write. Tools enforce it.

Section 05 · architecture

How a request becomes a customer outcome.

Three input surfaces converge through a single dispatch into the harness loop, and out to the platform's persistent stores.

user_chat chats domain customer message in
autopilot Publisher event fan-out
business_assist workspace UI operator action
DISPATCH harness.Dispatch resolve mode · actor · capabilities · budget

↻ THE HARNESS LOOP

load context · call LLM · execute tool calls · feed back · until done

max_steps = 10 · budget enforced · every step traced
Skills instruction store
Rules always-on
Memory explicit R/W
Tools built-in + MCP
LLM driver pattern
Trace append-only

Section 06 · context

What the agent sees on every call.

Eight layers, assembled fresh per run. The first four are cached on the model side; the last four vary every time.

01
System Identity who is the agent, who is the workspace
cached
02
Always-on Rules brand voice, PII policy, compliance
cached
03
Mode Guidance chat-mode etiquette vs autopilot tone
cached
04
Active Skills the operator's relevant playbooks
cached
05
Memory Snippets customer profile, conversation summary
per-run
06
Conversation recent thread (chat mode only)
per-run
07
Tool Definitions what the agent is allowed to call
per-run
08
Current Input the message or event that triggered this run
per-run

The Omazy harness builds the context fresh every time it speaks to the model. Static parts — workspace identity, rules, mode guidance, the active skills — sit on top and get cached server-side, so the model doesn't pay for them again.

Variable parts — memory snippets the agent chose to read, the live conversation, the tools it's allowed to invoke right now, and the input that triggered this run — go at the bottom. The contract is simple: nothing enters the prompt unless it's accounted for in one of these eight layers.

Why this matters

Prompt caching cuts first-token latency by ~40% on warm runs. Operators get faster replies, the platform pays a smaller bill, and every layer is auditable from the trace.

Section 07 · scale

One harness. Many workspaces.

A single multi-tenant runtime serves every customer. Their differences live in data, not infrastructure.

SHARED RUNTIME

harness.Runner

stateless · multi-tenant · one Go process

@acme Pro
skills24
tools11
langEN
@beti Free
skills8
tools5
langBN · EN
@orient Business
skills41
tools18
langEN · AR
@kiyora Pro
skills17
tools9
langEN
@matsuri Business
skills12
tools7
langJP
@voltura Enterprise
skills33
tools14
langEN · ES

Adding a new workspace is a database row, not a new server. Skills, rules, tools, memory, and budgets all live as per-tenant configuration loaded into the same runtime on demand.

Section 08 · governance

Rules over rules.

Guardrails layer top-down. Each tier appends to the one above it. A workspace can only override a platform rule with a logged compliance waiver.

01

Platform

Omazy-shipped — applies to every workspace. PII, profanity, factuality.

no-pii-leak.mdno-profanity.mdfactuality-vs-kb.md
02

Workspace

Workspace-authored — applies across all of one organisation's apps.

acme-brand-voice.md
03

App

App-scoped — applies only when running inside one app/brand.

acme-cards-tone.md
04

Mode

Per-mode tightening — chat mode might forbid what autopilot allows.

chat-no-billing-changes.md

Rules append, never silently override. The compliance trail is queryable: every waiver is one row in rule_waivers.

Section 09a · memory

Four memory layers.

Each layer has a purpose, a store, and a TTL.

Hot conversation Redis TTL · 24h

Current chat thread, in-flight scratchpad, intra-turn tool results.

Customer memory Postgres + pgvector TTL · durable

One row per customer profile. Preferences, sentiment, last-N session summaries.

Workspace memory Postgres TTL · durable

Org-level facts: open campaigns, holidays, pricing changes. Hand-authored or auto-written.

Trace store Postgres + ClickHouse TTL · 90d / inf

Full append-only run history. Replay, billing, future fine-tuning.

Section 09b · tools

Tool registry.

Built-in tools ship with the harness. Vendor tools arrive through the marketplace.

Chats

  • tag_session
  • assign_to
  • escalate_to_human
  • send_followup
  • attach_form_to_reply

Comms

  • send_email
  • send_sms
  • send_push

KB

  • kb_search
  • kb_read_doc

CRM

  • customer_lookup
  • customer_history

Memory

  • memory_read
  • memory_write
  • memory_search
  • memory_summarize

Workflow

  • schedule_followup
  • wait_for_event
  • sub_agent_call

Human

  • human_approval (suspend → operator click)

Shopify

  • lookup_order
  • create_refund
  • check_inventory

Stripe

  • get_invoice
  • refund_charge
  • list_subscriptions

Zendesk

  • create_ticket
  • find_similar
  • set_status

WhatsApp

  • send_template
  • open_session_window

Slack

  • post_message
  • lookup_channel

Section 10 · partner integration

AICoach adapter.

A two-way sync layer that lets workspaces share skill libraries and MCP credentials with aicoach.pw — the operator-coaching platform.

SOURCE OF TRUTH FOR PLAYBOOKS

  • · Skill packages (markdown + frontmatter)
  • · Operator coaching prompts
  • · Vendor MCP credentials (vaulted)
  • · Pre-tuned guardrails per industry

RUNTIME · MANY WORKSPACES

  • · workspace_skills (synced)
  • · workspace_tools / MCP (synced)
  • · trace + run history (telemetry back)
  • · per-workspace overrides preserved

The adapter sits between the two systems and treats AICoach as a content authority for generic playbooks (e.g., "how to handle a refund request" for an e-commerce vertical). Omazy workspaces inherit those skills automatically, layer their own brand-specific overrides on top, and stream anonymised trace data back so AICoach can improve the next version.

// workspaces opt-in per skill bundle. Overrides always win locally. No telemetry without consent.

Section 11 · the new chat flow

From "hello" to "thanks, that helped."

A complete customer chat, end to end.

  1. 01
    Customer sends a message Web widget · WhatsApp · email · SMS
  2. 02
    chats domain stores + emits an event EventSessionMessage in the catalogue
  3. 03
    Harness runs in user_chat mode Loads skills · rules · memory · tools
  4. 04
    LLM picks tool calls (or replies) kb_search → customer_lookup → reply
  5. 05
    Reply streams to the customer Text + interactive buttons / forms
  6. 06
    Customer clicks → harness resumes Same run thread, full audit trail

Section 12 · glossary of decisions

Ten decisions that hold the design together.

Each one is a non-obvious choice. If anyone disagrees, RFC 2203 is back open.

Instructions, not graphs
Operators describe intent in markdown. The LLM picks the order of tool calls. Drag-and-drop graphs are reserved for deterministic data pipelines.
One harness, three modes
Customer chat, business autopilot, and operator copilot are three configurations of the same underlying engine — not three codebases.
One process, per-tenant config
A single multi-tenant runtime serves every workspace. Adding a customer is rows in a database, not a new server.
Memory is explicit
The agent reads and writes memory through visible tools. Nothing leaks into the prompt invisibly. Auditable by default.
Tools are the security boundary
Permissions are enforced by code at the tool boundary, not by prompt instructions. The LLM cannot escalate by asking nicely.
RFC 2202 stays — v3 adds skills
The existing automations engine keeps running. Skills are an additive new layer; nothing is forced to migrate.
Driver pattern for the LLM
Anthropic primary, OpenAI fallback, on-prem optional. Swapping providers is a single file change, not a refactor.
Prompt caching, by design
Static parts of every call (rules, skills, persona) are cached on the model side. Cuts latency by ~40% on warm calls.
Append-only, full-fidelity trace
Every model call, tool call, and guardrail check is recorded. Foundation for replay, debugging, eval, and compliance.
Hooks borrowed from Claude Code
pre_run, pre_tool_call, post_reply. Workspaces extend the harness without forking the codebase.

// sources: docs/rfc2203-agentic-harness.md

// related: technical deep-dive →