Automation · one engine, three surfaces
The green results that lied.
A workflow engine is mostly a testing problem wearing an architecture costume. This is how a node-graph automation engine got built on top of an engine that already existed, and the four separate times a test came back green while proving nothing at all.
Section 01 · the expensive mistake we nearly made
Every automation feature wants to be a new engine. Almost none of them should be.
We already had two things that ran work in the background. A rule engine that fired a linear chain of actions on an event. A task engine that gave a model some instructions, a tool allowlist and a budget, then ran it durably with resume, approvals, scheduling and spend caps. The plan was to add a third: a visual graph engine with its own runs table, its own state machine, its own scheduler.
Reading the design against the code, roughly seventy percent of the proposed engine was a second implementation of the task engine. Not similar to it. The same eight run states, the same resume path, the same approvals, the same cron, the same budget ceiling. The genuinely new parts were branching, named ports, a compiler and a canvas.
The reframe that saved a quarter of work: a rule is a graph. It is literally trigger, then action, then action. A task is a graph too, a one-node one wrapping an agent loop. Once you see that, you are not building a third engine. You are building a compiler and a canvas in front of the engine you already have.
flowchart TB
subgraph BEFORE["BEFORE · three engines, one job"]
R1["rule engine
trigger to action chain"] --> D1[("runs table A")]
R2["task engine
instructions plus tools"] --> D2[("runs table B")]
R3["graph engine
proposed"] -.-> D3[("runs table C")]
end
subgraph AFTER["AFTER · one engine, three surfaces"]
S1["rule
two fields"] --> CIR["one graph format"]
S2["task
one agent node"] --> CIR
S3["graph
explicit branches"] --> CIR
CIR --> COMP["one compiler"]
COMP --> RUN["one runner"]
RUN --> DB[("one runs table")]
end
The practical test for this in any codebase: write down the columns each system needs and lay them side by side. If the middle column already has most of what the new one wants, you are not designing a system, you are duplicating one. The cost of duplicating shows up later as two of everything, permanently.
Section 02 · the decision everything else hangs off
A third of the node types do not finish. The interface has to admit that.
The obvious interface for a node is a synchronous execute that takes an input and returns an output. It is also wrong, and it is wrong in a way that is very expensive to discover after the runner is written.
Count the node kinds that cannot return synchronously. An approval waits on a person, possibly for days. A delay waits on a timer. A trigger never executes at all, it binds. Anything dispatched to a remote worker or a sandbox hands off now and gets a result later. That was eight of our twenty-odd kinds, plus two whole runtime classes.
Force those through a synchronous call and you get one of two bad outcomes. Either a goroutine parks per in-flight node, each holding a span and probably a transaction, which is how a hundred parked approvals become an outage. Or execute returns a sentinel while the real result arrives somewhere else, which is an interface that lies to every caller.
So the return type is a sum: Done | Pending. Done carries the output. Pending carries an opaque continuation token, a reason, and a mandatory deadline. The control plane persists the token, frees the goroutine, and calls resolve when the outside world produces an answer.
stateDiagram-v2 [*] --> Dispatched: control plane checkpoints, then enqueues Dispatched --> Running: worker leases, fence increments Running --> Done: Start returns Done Running --> Pending: Start returns Pending with a deadline Pending --> Resolving: person answers, timer fires, worker calls back Resolving --> Done: Resolve returns Done Pending --> Failed: deadline elapsed Running --> Failed: error, policy is fail Done --> [*] Failed --> [*]
Costed honestly: this shape takes an afternoon to design up front. Retrofitting it after the runner and the worker pool exist is a rewrite of every node. When an interface has to express suspension, decide that on day one.
Section 03 · the ordering that survives a crash
Write the row, then do the work. Never the other way round.
Durability in a run engine is mostly one rule about ordering, and the rule is easy to state and easy to get backwards.
Every node is checkpointed to its own row before the work is dispatched. A node row with no result is a node to retry, which is recoverable. A dispatch with no row is work that happened and nothing knows about, which is not. The asymmetry is the whole point: one direction leaves you a puzzle you can solve, the other leaves you a side effect you cannot find.
Per-node state goes in its own table rather than a JSON column on the run. Postgres has no partial JSON update, so a document rewrite costs the whole row, the TOAST chain and a full write-ahead record on every transition. Worse, two branches finishing at once take the same row lock, which turns a graph engine into one with a global mutex per run.
Then the harder half. Delivery is at least once. A worker can be alive but slow, hit a garbage collection pause, have its lease expire, and a second attempt starts while the first is still running. Both then try to report. Deduplicating the result is not enough, because you can discard a duplicate result but you cannot discard a sent email.
sequenceDiagram participant C as control plane participant A as worker A participant B as worker B participant V as vendor C->>C: write node row, fence = 1 C->>A: dispatch A->>V: send Note over A: alive but slow C->>C: lease expires, fence = 2 C->>B: dispatch attempt 2 B->>V: send with same idempotency key V-->>B: deduplicated, one effect A-->>C: result carrying fence 1 C-->>A: rejected, stale fence B-->>C: result carrying fence 2 C-->>B: accepted
Made structural, not conventional: any capability declaring that it sends something must also declare how it avoids sending twice. That is a required field on the descriptor, checked when the capability registers at boot, and the compiler refuses to publish a graph containing a sending node without one. A code review convention would have been forgotten within a month.
Section 04 · the guardrail that starves you
The single-flight rule that is right for a nightly job silently drops customer events.
Scheduled work has an obvious concurrency rule: do not run the nightly digest twice. A partial unique index over the non-terminal states enforces it in one line, and the second starter loses the insert race atomically. For a scheduled task that is exactly right.
Lift the same index onto event-triggered work and it starves. Here is the worked example, using our own design document's own example graph. A workflow fires when a conversation closes, and it pauses at a manager approval. The first conversation closes at nine in the morning and parks waiting for a manager who is in a meeting. With one non-terminal run allowed per workflow, and an overlap policy that defaults to skipping, every conversation that closes for the next forty minutes is silently dropped.
Two corrections. The key is per entity, not per workflow, so different conversations run beside each other and only the same conversation queues behind itself. And a suspended run releases the slot entirely, because a run waiting on a person is not a run occupying a worker. The index predicate names only the two states that genuinely hold something.
| Trigger | Sensible key | What the wrong key does |
|---|---|---|
| Schedule | Empty, so one run at a time overall | Nothing. This is the correct default here. |
| Manual | Empty | Nothing. Every run is someone deliberately pressing a button. |
| Event, per conversation | The conversation id, compiled from the trigger as {{ trigger.session_id }} | Keyed on the workflow, one parked approval drops every other conversation for as long as it waits. |
This is also a product control wearing a storage costume, so it is exposed to the author in plain language on the trigger, not buried. "One run at a time per conversation" against "one run at a time for the whole workflow" is a business decision, and it belongs in a dropdown with a sentence next to it.
Section 05 · the launch blocker nobody names
The first time someone presses Run, it must not be able to email a real customer.
Every design review of an automation builder covers the engine, the compiler and the canvas. The thing that actually generates the unrecoverable support ticket rarely appears on the list.
A draft workflow runs against live credentials. Someone learning the product builds a refund confirmation, presses Run to see what happens, and the product emails forty real customers. That ticket does not get fixed by a patch, because the emails already went.
So a run carries a mode, and in test mode every node that reaches outside the team stubs itself: it records what it would have sent and returns that as its output. Everything else runs for real, including reads, because stubbing a read teaches nobody anything and the whole point of a test run is to see real data flow through the graph. The console shows the recorded payload, so the operator reads the message that did not go out.
Two details that make it hold. A test run is a distinct kind of run, so it appears in the run list with full history and is excluded from the billing counter, because charging people to learn your product is hostile. And a timer node does not actually wait in a test run, because watching a run sit for six hours teaches nobody whether the graph is correct.
Verified rather than assumed: the production smoke test asserts on the recorded stub, not on a green run. A test run that quietly did nothing at all would also finish successfully.
Section 06 · the smallest language that stays small
Two decisions keep an expression syntax from growing into a language by accident.
Values move between steps through a narrow substitution syntax: path reads against the trigger payload, an earlier step's output, and a registry of values resolved fresh at fire time such as today or yesterday.
The first decision is what a missing path does. Write {{ nodes.n3.output.name }} where that step sat on a branch that was never taken, and something has to happen. The best-known tool in this category fails the node, and it is the single most-cited authoring complaint against it, because a value that was simply absent takes down a run far from where it went missing. A missing path resolves to null instead, and null is only an error where the target schema marks the field required.
The second is the filter list, which is closed at five entries and is not an extension point. A default filter earns its place because without it every trivial graph pays a tax: you needed one string fallback, so you added a code node. That is how expression syntaxes grow into languages nobody chose to build.
A dynamic value also has to be recomputed on every firing rather than frozen when the workflow was authored. "Yesterday" has to mean yesterday on every run, which is exactly why a literal cannot express it and why the registry exists.
One trap worth naming: the idempotency key for a scheduled firing is derived from the schedule slot, never from the resolved inputs. Include a dynamic value like today in the key and every firing becomes unique, which defeats the deduplication the key exists for.
Section 07 · each layer earns its place
Four layers of testing, and the specific bug only that layer could find.
A test suite justifies itself by the class of defect it catches that nothing cheaper would have. Here is what each layer actually found, rather than what it was supposed to find.
flowchart LR U["unit tests
compiler rules"] -->|"caught"| U1["a graph that
loops forever"] I["integration tests
real database"] -->|"caught"| I1["a column that
cannot scan null"] P["production smoke
real deploy"] -->|"caught"| P1["a schedule that
never fired"] B["browser tests
touch emulation"] -->|"caught"| B1["a checker that
tested nothing"]
Unit: one violating graph per rule
The compiler has around fifteen validation rules. Each one gets a graph that breaks it: two nodes sharing an id, an edge leaving a port the node does not have, a cycle, an expression reading outside the allowed roots, a capability nothing implements. A compiler whose only test is the happy path is a compiler that publishes anything.
Integration: hand-written SQL against a real server
Pure logic tests cannot see a column mismatch. The first run against a real database failed immediately: a raw JSON type cannot scan a null, so every read of a workflow that had never been published would have returned a server error. That bug is invisible to any amount of mocking.
Production: exercise the path you changed
A green container proves a process started. The post-deploy script created a workflow from a template, published it, ran it, and asserted on the recorded stub. Then it armed a schedule and waited to see whether anything fired.
Browser: layout properties a screenshot cannot assert
Two machine-checkable things: no horizontal overflow at a phone width, and no interactive control smaller than the minimum touch target. Both are tedious for a human to check on every view and trivial for a browser to measure.
Section 08 · the part worth reading twice
Four times a test came back green while proving nothing.
The engine work was ordinary. The genuinely instructive part of this project was how often a passing result meant nothing at all, and how each one was caught. A false green is worse than a red, because a red gets investigated.
1 · A test that passes either way
A regression test was written for a bug found in production. It passed. Then it was run again with the fix reverted, and it still passed, which made it decoration. It needed one more assertion before it reproduced the actual symptom. A regression test you have not seen fail is not yet a regression test.
2 · An empty page passes every layout check
The browser suite reported a clean pass across every viewport. The screenshots showed a loading spinner, and then an error card. A page with nothing on it has no horizontal overflow and no small tap targets. Assert that the thing under test actually rendered before you assert anything about it.
3 · A narrow desktop is not a phone
The mobile checks ran a desktop browser resized to phone width. That reports a fine pointer, so every touch-specific rule stayed inert and the suite was measuring something no user ever sees. It now runs in a touch context and throws if the coarse pointer media query does not match, rather than quietly measuring the wrong thing.
4 · Unchanged code type-checks perfectly
Two scripted edits silently failed to match their target text. Type-check passed, lint passed, the build passed, and the change was committed. It had never been applied. An edit tool that does not fail loudly on a miss will eventually let you ship nothing and call it done.
The pattern underneath all four: each check was measuring a proxy rather than the property. Not "does the UI lay out correctly" but "does this DOM have overflow", which an empty DOM satisfies. Not "was the fix applied" but "does the code compile", which unchanged code satisfies. The habit that catches all of them is the same: before trusting a green, ask what a false pass would look like, and make the check able to tell the difference.
Section 09 · non-fatal, and therefore silent
The workflow published successfully. It just never ran.
The best bug of the project only appeared against a real production database, and it is a good illustration of how a defensive design choice can hide the thing it was protecting.
Publishing a workflow writes its schedule out of the trigger. That write is deliberately non-fatal: if binding fails, the publish still succeeds and the failure is logged, because losing a publish over a scheduling detail would be worse. Sound reasoning, and it turned a hard failure into a silent one.
The statement set a timestamp column through a conditional whose other branch was an untyped null, roughly CASE WHEN enabled THEN $4 ELSE NULL END. Postgres inferred the parameter type from that expression, landed on text, and rejected the whole update. The publish returned success. The schedule column stayed empty. The workflow then sat there forever, looking perfectly healthy, firing nothing.
The fix is one explicit cast, CASE WHEN enabled THEN $4::timestamptz ELSE NULL END. The interesting part is why nothing caught it earlier. There was a unit test, and it tested the function that reads the trigger and produces the binding. It stopped exactly at the seam. The bug lived one layer further down, in a statement only a real database would ever reject.
Two things changed afterwards. The regression test asserts on the stored columns after a publish, not on the binding function's return value. And the post-deploy check for anything that fires on its own now waits to observe a real firing, because "the schedule was saved" and "the schedule fires" are different claims.
Section 10 · one schema, four surfaces
The palette, the settings panel, the tool definition and the assistant all read the same declaration.
Every node kind ships a descriptor: a title, a description written for a model as much as a person, schemas for its settings, input and output, its side-effect class, and what it needs to run.
That one declaration drives the canvas palette, the settings panel, the definition handed to a model that can call it as a tool, and the assistant that drafts graphs. It only stays true if the frontend never learns a node kind by name. Ours renders every field from the schema plus a small set of interface hints, so adding a capability on the server puts it on the canvas with no frontend release at all. An unknown hint degrades to a plain text input rather than failing.
There is a discipline that goes with this, learned from an earlier mistake in the same codebase. A console once advertised nineteen tools while the backend implemented six. The other thirteen were silently dropped at run time, so anything authored against them simply never did that part of its job, quietly, on every run. On a visual palette the same gap is far more expensive. So an unregistered capability is now a hard compile error rather than a skipped entry, and a test fails the build if the console offers anything the backend cannot execute.
On mobile, the canvas is not the job. Authoring a graph on a phone is a bad idea, but reading one, triaging a run and answering an approval all happen on a phone constantly. The first mobile layout stacked the palette on top, which put a list of node names above the fold on the one device where authoring is not what anyone is doing, and pushed the graph off screen entirely.
Section 11 · what generalises
Six things worth taking to a different codebase.
Compare the columns before you build
Put the fields your new system needs beside the fields the existing one has. Heavy overlap means you are duplicating, not designing. Two of everything is a cost you pay forever.
Decide suspension on day one
If any unit of work waits on a person, a timer or another machine, the return type has to say so from the start. Adding it later is a rewrite of every implementation.
Checkpoint before you act
A record with no work is recoverable. Work with no record is not. Under at-least-once delivery, pair a fence for results with a vendor idempotency key for effects, and make declaring one a requirement of the type rather than a convention.
Give people a safe way to try it
Anything that can reach a customer needs a mode that provably cannot, and it should be the default on the button someone presses first. Show them what would have been sent.
Fail loudly at the boundary
Silent skips, non-fatal writes and tolerant parsers all trade a loud failure now for a quiet wrong answer later. Where you genuinely need one, pair it with a check that would notice.
Ask what a false pass looks like
Before trusting a green result, describe the failure that would still produce it. If you cannot tell those apart from the output, the check is not finished yet.
Section 12 · shape of the thing
What shipped.
A graph format, a compiler, a ready-set runner on the existing durable engine, a capability registry, triggers, and a canvas. The run engine underneath is the one that was already there.
Engine
- Graph format with immutable node ids and named ports, so a rename is a display change and a diff reads as English
- Compiler producing an execution plan: port map, join semantics, cycle rejection, declared side-effect set, cost estimate, content hash
- Ready-set walk rather than a topological sort, because merges and future loops break a linear order
- Per-node rows with fences; suspension through a continuation token with a mandatory deadline
- Immutable published versions, so a run suspended across a publish resumes on the plan it started with
Around it
- Capability registry with an adapter for existing tools, so no subsystem had to be rewritten to join
- Schedule and event triggers, bound at publish rather than at edit, so a half-finished draft cannot fire
- Throttle shipped with the triggers: a stop switch, a per-tenant pause, and a daily ceiling denominated in money rather than run count
- Canvas, palette and settings panel rendered entirely from the capability declarations
- A seeded gallery, because the first session on a blank canvas usually ends without a run
The throttle landing beside the triggers rather than a few milestones later was deliberate. The original sequencing had unattended, model-invoking, customer-facing automation running for several milestones with no ceiling and no stop button. Guardrails are cheap. The incident they prevent is not.