Definition · agentic harness
A loop that calls an LLM with tool definitions, executes the tools the LLM picks, and feeds results back until the LLM emits a final answer. Same shape as Claude Code, OpenAI Agents SDK, Cline.
Why one harness, three modes
user_chat / autopilot / business_assist all share the SAME loop. They differ in which skills are loaded, which tools are exposed, which guardrails apply, whose identity is on the wire. Mode is config — not a separate codebase. Cuts duplicated tool-handler/trace/budget plumbing by 3×.
Why instructions, not graphs
Skill = markdown + YAML frontmatter ("when VIP session opens, greet warmly + route + ping #vip"). LLM reads it, picks tool order. n8n/Flowise solve a 2020 problem (operators couldn't code) — in 2026 operators describe intent in English and an LLM picks tools. Graphs are for ETL, not CX.
Engineering decision · explicit memory R/W
Memory is read/written through tools (memory_read, memory_write), never silently injected into the prompt. Auditability: trace shows exactly which memories shaped the output. Bounded context: LLM picks what it needs, not an opinionated retriever stuffing 8K tokens.
Engineering decision · capability boundary
Tools — not the LLM — enforce permissions. Even if the LLM "tries" to call a forbidden tool, the registry rejects and returns a structured error. Capabilities are the security perimeter, not prompt instructions.
Engineering decision · max_steps budget
Default 10 LLM-calls per run. Pre-tool hook rejects redundant calls (same args twice). Trace surfaces step explosion. Stops runaway loops cold.