Skip to main content
Omazy Engineering

Copilot · one task layer, every surface

Adding a Copilot capability is a declaration.

A workspace assistant that can actually operate the account has to change real settings, launch channels, and load knowledge, without becoming a second, more powerful way to bypass permissions. The design keeps that from happening by construction: every capability is a thin wrapper over an API the operator could already call by hand, so it inherits their exact permissions, and no write ever runs without a human confirm.

// aitask registry · /ai/tools · acts-as-current-user · confirm gate // RFC 2209 · registry (C0) shipped on main · deployed

Section 01 · the problem

Every feature was growing its own AI button.

The console kept sprouting one-off assistants: a Draft-with-AI on the test-suite screen, another on the agent editor, each wired straight to a model call. Useful in isolation, but there was no shared way for an operator to just say what they want and have the account change, and no single answer to the question that decides whether that is safe: when the AI acts, whose permissions is it acting with?

The tempting fix is a powerful assistant with its own service key that can do anything. That is also the dangerous one. It becomes a second control plane, more capable than the UI and gated differently, and the day its prompt is talked into something it should not do, it has more reach than the person who talked to it. We wanted the opposite: an assistant that can do exactly what the operator in front of it could do by hand, and not one thing more.

Section 02 · one idea

Two surfaces, one core, no new privileged path.

Draft-with-AI and the chat Copilot are the same engine. One fills the form you are on; the other holds a conversation and runs account tasks. Both resolve the caller's identity from the request they already arrived in, and act as that person.

flowchart TB
  A["Draft with AI
(single-shot, autofill)"] --> CORE B["Copilot chat
(bounded tool loop)"] --> CORE subgraph CORE["AI Task core · runs AS the current user"] G0["preflight: can(actor) · rate-limit · meter"] PR["prompt: GLOBAL > TASK > tenant > context(DATA)"] EX["A: forced tool-call → stage fields
B: reads inline · writes → confirm-card"] end CORE --> LLM["internal/llm Router → gateway"] CORE --> H["internal/harness loop + action tools"]
Fig 1. Surface A forces a single schema-constrained tool call to stage form fields. Surface B runs a bounded tool loop. They share the gateway, the prompt hierarchy, the task registry, the permission gate, and the audit. Reads execute inline; writes come back as a confirm-card.

A capability is not a bespoke integration. It is an AI Task: a thin wrapper over an endpoint the operator could already reach. The model never gets a private door into the database. It gets the same front door the UI uses, opened inside the operator's own authenticated request.

Section 03 · the load-bearing rule

It acts as you, and writes wait for you.

The permission model is not new code. The Copilot's tools run inside the operator's request, with the same role and workspace already resolved by the auth middleware, so a member's Copilot is exactly as constrained as a member.

flowchart LR
  ASK["operator: 'add /fleet to knowledge'"] --> PROP["model proposes kb.add_url"]
  PROP --> CARD["confirm-card
(nothing runs yet)"] CARD --> OK{"operator confirms?"} OK -->|"yes"| RUN["runs the SAME service the UI button uses
behind RequireWorkspaceRole"] OK -->|"no"| STOP["discarded · audited"]
Fig 2. A write is proposed, not performed. The operator sees the exact call before anything happens. On confirm it runs the same service the UI button calls, behind the same role gate, and the held call executes verbatim, so the model is never re-asked what it wanted to do.

Two properties fall out of this for free. Permissions are correct because they are the operator's own, checked when the tool is offered and again when it runs, so a stale card cannot escalate. And nothing irreversible is a surprise, because a write is a proposal until a human approves it. The confirm-and-run half reuses the runbooks harness, where holding a side-effecting call and later executing the persisted version verbatim is already how the approval gate works.

The principle: the cheapest way to give an AI the right permissions is to make it borrow a human's. There is no new authorization surface to get wrong, because there is no new authorization surface at all.

Section 04 · autofill, typed

The form is the schema.

Draft-with-AI does not parse free text into a guessed object. The task names a form, and that form's field definitions are the contract the model must fill.

Each form is a list of typed controls. The server turns those controls into a JSON tool-schema, one property per field, with the field's type and rules carried across, then forces the model to answer in that shape. A colour field becomes a hex string, a dropdown becomes an enum of its options, a repeater becomes an array of nested objects. The reply is validated against that schema, and any field the model cannot ground in real account data is abstained rather than invented.

// The form IS the schema. No hand-written object per task.
// Control[] (rfc2206) → generated JSON tool-schema → forced tool_choice → validated:
type AutofillResult = {
  fields:    Record<ControlKey, unknown>        // keyed to the form's controls
  perField:  Record<ControlKey, { source }>     // provenance per field
  abstained: ControlKey[]                        // couldn't ground → left blank
}
Fig 3. There is no hand-written object per task. The form defines the fields, the server derives the schema, and the result carries per-field provenance plus an explicit list of what it left blank. The staged values fill the form for review; the operator's Save is the authorization boundary.

This is why one autofill engine serves every form. Add a new admin form and it can be drafted by AI the day it ships, with no per-form model wiring, because the thing that types the fill is the form the operator already sees.

Section 05 · what shipped

One declaration, and the Copilot can do it.

The first piece on main is the capability registry: a coded catalog that is the single source of truth for what the Copilot can do, and the gate for who may do it.

// internal/aitask/registry.go — one entry, and the Copilot can do it.
{
  Key:         "widget.personalize",
  Kind:        KindAction,          // text | autofill | action
  SideEffect:  SideEffectWrites,    // writes/sends are held for confirm
  MinRole:     "manager",           // from the wrapped route's real gate
  Permission:  "settings.edit",     // rbac key, checked at propose AND confirm
  Entitlement: "widget",            // hidden if the plan lacks it
  TargetForm:  "widget.appearance", // autofill: this form's Control[] types the fill
  Wraps:       "PATCH /workspace/:h/apps/:a/settings/widget",
}
Fig 4. A capability declares its kind, its side-effect, the role and permission taken straight from the gate on the endpoint it wraps, an optional entitlement, and, for autofill, the form that types the fill. That is the whole registration.

A single endpoint, GET /workspace/:handle/ai/tools, returns the tasks the caller may run, already filtered by their resolved role and grouped for the prompt guide. It replaces the hardcoded lists that each screen used to carry, so the UI and, next, the chat dispatcher read the same catalog. A member asking for the list simply does not see a manager-only task like widget.personalize; they do see the member-level ones like kb.add_url.

Role at the door

Each capability's min_role and permission are lifted from the real gate on the route it wraps, so the catalog can never advertise a task the caller could not perform by hand.

Plan-aware

An entitlement field lets a task hide when the plan lacks it, so the Copilot offers WhatsApp only where it is available, and can suggest an upgrade instead of failing.

Section 06 · what it unlocks for developers

Ship the endpoint, get the AI for free.

The point of a task layer is that the hard parts are solved once. A team adding a capability does not build an assistant. They declare a task, and inherit everything the core already provides.

You bring

  • An endpoint that already exists and is already permission-gated.
  • One registry entry: key, kind, side-effect, the gate, an optional form.
  • For an action, a thin tool that calls that same service.

You get

  • A prompt-guide entry and a place in /ai/tools, filtered by role and plan.
  • Draft-with-AI on the form, typed by the form, with no model wiring.
  • The confirm gate, verbatim execution, and the audit trail, unchanged.

The larger idea is that an internal AI agent for the account does not need special powers to be useful. It needs the account's own doors, opened as the person who asked, with a pause before anything is written. What that leaves for us to build is not a smarter, riskier assistant. It is a longer, ordinary list of endpoints, each one a task the Copilot can now run on your behalf.

Sources

  • docs/rfc2209-ai-task-copilot.md — the two-surface, user-scoped design.
  • internal/aitask/ — the capability registry and /ai/tools (C0).
  • internal/admin/ai — the AI Manager: ai.run and the versioned task registry.
  • internal/harness — the tool loop and the verbatim confirm gate the Copilot reuses.
  • internal/rbac, src/forms — the permission matrix and the form contract that types autofill.