Webhooks · signed event delivery
Your events, pushed and signed.
A workspace should not have to poll us to learn that a chat session opened or an agent went live. Omazy pushes each domain event to any HTTPS endpoint you register, as a stable JSON envelope carrying an HMAC-SHA256 signature over the exact bytes we sent. Here is how it is built, what lands on your server, how the engine backs off when you rate-limit it, when it gives up on a dead endpoint, and how to verify a delivery in about fifteen lines.
Section 01 · the problem
Polling is a tax on everyone who integrates with you.
Every workspace that wants to react to Omazy in their own systems, a CRM row on a new lead, a Slack ping when an agent publishes, a spreadsheet of closed sessions, faces the same choice: poll our API on a timer, or wait for us to push.
Polling is the worse deal on both sides. The integrator writes cursor bookkeeping and eats latency equal to their poll interval. We serve a constant baseline of requests that mostly return nothing. And the interesting events, the ones worth reacting to, are rare enough that a one-minute poll is 99 percent waste and still a minute late.
So webhooks. The design constraint that shaped everything else: a pushed event has to be provably ours. Anyone on the internet can POST JSON at your endpoint claiming a session just closed. Without a signature, a webhook receiver is a public, unauthenticated write path into your business logic. That is why the signature is not a feature of this system; it is the system.
Section 02 · architecture
One event bus. Two consumers. No second vocabulary.
Workspace automations already listened to a stream of domain events. Rather than build a parallel emitter for webhooks, the webhooks service registers as a second sink on the same publisher. An event is emitted once and reaches both an in-platform automation and an out-of-platform endpoint.
flowchart LR
subgraph CORE["domain"]
EV["session.opened
agent.published
kb.updated · …"]
end
EV --> PUB["automations.Publisher"]
PUB --> AUT["workspace automations
(in-platform reaction)"]
PUB -->|"EventSink"| WH["webhooks.DispatchEvent
(out-of-platform push)"]
WH --> Q{"active subs
for this event?"}
Q -->|"none"| STOP["return, no work"]
Q -->|"n"| GO["n goroutines · one POST each"]
GO --> EP["your HTTPS endpoint"]
GO -.-> LOG[("webhook_deliveries")]
automations.Publisher fans each domain event to in-platform automations and, through an EventSink adapter, to webhooks.DispatchEvent. Every active subscription gets its own goroutine, so one slow receiver never delays another. Each attempt is recorded whether it succeeded or not.Why one vocabulary, not two
The webhook event catalog deliberately mirrors the automations catalog. A workspace that learns session.closed to build an automation already knows the webhook event name, and there is no drift where an automation fires but a webhook does not. The catalog is redeclared in the webhooks package rather than imported, so the subsystem validates subscriptions against its own surface and the console can render the identical checklist.
Why detached goroutines
Dispatch returns immediately. Each delivery runs on a background context bounded by the dispatch timeout, not the publishing request's context, because the request that caused the event has no business waiting on your server. The bound is what stops a hung receiver from leaking a goroutine forever.
Section 03 · the event registry
Where the event names come from.
There is no webhook-specific event vocabulary. The platform's domain-event registry is a single closed set of constants in the automations package, and the webhook engine subscribes against that same set. What a workspace can trigger an automation on is exactly what it can receive a webhook for.
flowchart LR REG["internal/automations/registry.go
EventKind · 6 domain events"] REG -->|"producers call Publish"| PUB["automations.Publisher"] PUB --> A["automations engine"] PUB --> WH["webhooks.DispatchEvent"] REG -.->|"redeclared, not imported"| WREG["internal/webhooks/registry.go
same 6 + webhook.test"] WREG -->|"validates subscriptions"| WH subgraph OTHER["separate vocabularies · deliberately not on this bus"] TEL["webingest · browser telemetry"] RT["realtime · WebSocket room events"] end
Redeclared, not imported
The webhooks catalog is a copy of the automations catalog plus webhook.test. The upside is that the subsystem validates subscriptions against its own surface and can expose the list the console renders without an import cycle through the automations engine. The cost is honest: the two lists are joined by discipline, not by the compiler, so adding an event in one place and not the other fails silently. That is a known seam, and the right fix when it starts to hurt is a shared package both import rather than a third copy.
What is deliberately not on this bus
Two other things in the platform look like events and are not. Browser telemetry from the widget (internal/webingest) has its own allow-list of client-emitted names, is fire-and-forget, and is high volume by design. WebSocket room events (internal/realtime) are transport-level fan-out to a live UI. Neither is a durable business fact, and putting either on the webhook bus would turn a customer's endpoint into a firehose.
Status, plainly stated: the registry, the publisher, the sink, and the delivery engine are all shipped and wired. The producer call sites are the remaining step: the domains that own these facts do not yet call Publisher.Publish in their hot paths, which the automations package documents as a follow-up. Until they do, the ping is the only thing that reaches a registered endpoint, and any automation on the same six events is idle for the same reason. Because both consumers hang off one publisher, wiring a producer once lights up both at the same time.
Section 04 · the envelope
Four fields, the same shape every time.
Every delivery is a flat JSON object with four keys: the event name, the workspace it belongs to, a per-event data payload, and the send timestamp. The test ping uses the identical shape, so your parser exercises the same code path for a ping as for a real event.
POST https://your-endpoint.example/hooks
Content-Type: application/json
User-Agent: Omazy-Webhooks/1.0
X-Omazy-Signature: sha256=52c9141de7b8cd0e30f0604ddbee1df47e54ef7f27f9fdac377086691fa0d4c4
{"event":"webhook.test","workspace_id":"0e4b758f-39a0-4923-bc5e-2b6524add7a1","data":{"message":"This is a test ping from Omazy."},"sent_at":"2026-07-31T01:45:35.404112686Z"} sent_at carries nanosecond precision because Go writes RFC 3339 that way. Treat it as an opaque string when verifying. Any parse-and-reformat round trip will change the bytes and break the signature.
The subscribable catalog
| Event | Fires when |
|---|---|
session.opened | A new chat session began. |
session.closed | A chat session was resolved or auto-closed. |
agent.published | An agent draft was promoted to active. |
kb.updated | Knowledge-base material changed. |
member.joined | A new workspace member landed. |
budget.alert | A workspace LLM-budget threshold was crossed. |
There is a seventh kind, webhook.test, that is intentionally not subscribable. It is emitted only by the explicit ping button, and the ping fires regardless of which events a webhook is subscribed to, so you can validate a brand-new endpoint before wiring it to anything real. A ping is also honoured on an inactive webhook, because "does my endpoint work" is a question you ask before you turn it on.
Section 05 · signing
HMAC over the exact bytes on the wire.
The envelope is marshalled once. That byte slice is what gets signed and what gets sent, in that order, with nothing in between. The signature travels in X-Omazy-Signature as the algorithm name, an equals sign, and lowercase hex.
// internal/webhooks/service.go
// sign returns "sha256=<hex>" where hex is HMAC-SHA256(secret, body).
func sign(secret string, body []byte) string {
mac := hmac.New(sha256.New, []byte(secret))
mac.Write(body)
return "sha256=" + hex.EncodeToString(mac.Sum(nil))
}
// ...and at delivery time, over the SAME bytes that go on the wire:
req.Header.Set("Content-Type", "application/json")
req.Header.Set("User-Agent", "Omazy-Webhooks/1.0")
req.Header.Set("X-Omazy-Signature", sign(w.Secret, body)) flowchart TB A["envelope marshalled once
body := json.Marshal(payload)"] --> B["sig = HMAC-SHA256(secret, body)"] B --> C["X-Omazy-Signature: sha256=hex"] C --> D["POST · same body bytes"] D --> E["your server reads the RAW body"] E --> F["recompute HMAC with your copy of the secret"] F --> G{"constant-time equal?"} G -->|"no"| H["401 · drop it"] G -->|"yes"| I["trust the payload, do the work"]
The whsec_ prefix is part of the key
A secret looks like whsec_ followed by 64 hex characters. The prefix exists so the string is instantly recognisable in a log, an env file, or a support ticket. It is not a scheme marker to strip, and the hex tail is not meant to be decoded to bytes. The HMAC key is all 70 characters as UTF-8. This trips up almost everyone once, which is why the test vector below prints what you get if you strip or decode it.
The secret is shown exactly once
It is minted in the console when you create the webhook, sent to the backend on create, and stored. The API never serialises it back: the model tags it json:"-" and returns only secret_last_4, so an operator can tell two webhooks apart without the full value ever leaving the server. Copy it at creation. There is no read-back and no rotate endpoint yet, so recovering a lost secret means delete and recreate.
Section 06 · verifying, for real
Fifteen lines, three rules.
Recompute the HMAC with your copy of the secret, over the raw request body, and compare the whole header value in constant time. Everything that goes wrong is a violation of one of those three.
| Rule | Why it bites |
|---|---|
Key is the full whsec_… string | Stripping the prefix or hex-decoding the tail produces a well-formed signature that never matches. No error, just a permanent mismatch. |
| Hash the raw body bytes | Re-serialising a parsed object is not the same bytes. Python's default json.dumps inserts spaces after separators. A Go map round trip sorts keys. Both silently change the hash. |
| Compare the whole header | The header includes the sha256= prefix. Compare it as one string, in constant time, so the check cannot be turned into a timing oracle. |
Pipedream, copy-paste
Paste this as a code step. Pipedream renders a masked input on the step where the secret goes, which avoids the most common first failure (see section 06).
import crypto from "crypto";
export default defineComponent({
props: {
secret: {
type: "string",
label: "Omazy webhook signing secret",
description: "The full whsec_... value, prefix included",
secret: true,
},
},
async run({ steps }) {
const secret = this.secret;
if (!secret) throw new Error("signing secret is not set on this step");
const event = steps.trigger.event;
// Prefer the raw body. JSON.stringify of the parsed object is a fallback,
// not a guarantee: it only matches while key order and separators survive.
const raw =
typeof event.bodyRaw === "string" && event.bodyRaw.length
? event.bodyRaw
: JSON.stringify(event.body);
const expected =
"sha256=" +
crypto.createHmac("sha256", secret).update(raw, "utf8").digest("hex");
const got = event.headers["x-omazy-signature"] ?? "";
const ok =
expected.length === got.length &&
crypto.timingSafeEqual(Buffer.from(expected), Buffer.from(got));
if (!ok) {
throw new Error(
`bad signature
expected: ${expected}
received: ${got}
usedRawBody: ${event.bodyRaw != null}
bodyLen: ${Buffer.byteLength(raw)} (compare to content-length)`
);
}
return { verified: true, event: event.body.event };
},
}); Note the error message. It prints both signatures plus the body length, so the moment it does fail you can compare bodyLen against the content-length header and immediately tell whether the key is wrong or the body reconstruction is.
Go and Python
func verify(secret string, body []byte, header string) bool {
mac := hmac.New(sha256.New, []byte(secret))
mac.Write(body)
want := "sha256=" + hex.EncodeToString(mac.Sum(nil))
return hmac.Equal([]byte(want), []byte(header))
} import hmac, hashlib
def verify(secret: str, body: bytes, header: str) -> bool:
want = "sha256=" + hmac.new(secret.encode(), body, hashlib.sha256).hexdigest()
return hmac.compare_digest(want, header)
# Flask: use request.get_data(), NOT request.json
# FastAPI: use await request.body(), NOT the parsed model In every framework, the trap is the same: the convenient accessor gives you the parsed body. Reach for the raw one. In Express that means express.raw() or capturing req.rawBody in the JSON parser's verify hook, and in Next.js route handlers it means await req.text() before any json() call, because the body stream can only be read once.
A test vector you can run offline
Before pointing anything at a live endpoint, check your implementation against a fixed triple. The secret here is a throwaway, not a real one.
secret whsec_4a1f0c9b7e2d5a8c3f6b9d0e1a2c4b6d8e0f1a3c5b7d9e0f2a4c6b8d0e1f3a5c
body {"event":"webhook.test","workspace_id":"0e4b758f-39a0-4923-bc5e-2b6524add7a1",
"data":{"message":"This is a test ping from Omazy."},
"sent_at":"2026-07-31T01:45:35.404112686Z"}
(174 bytes, no whitespace between tokens)
signature sha256=4c49aedadf7de2875d9e448a818bf51fb39c59a632a8b3ff15c9e01c574621a3
If your code returns 355cf1ee… you stripped the whsec_ prefix.
If it returns df763d6b… you hex-decoded the secret instead of using the string. Section 07 · two failures worth naming
Both were on the receiving side. Both looked like our bug.
Every integration we have watched hit one of these two. Neither is exotic, and neither produces an error message that points at the actual cause.
"The signature does not match the one I was given"
Two things wearing the word "signature". The whsec_… value is the key and never appears on the wire. The header is a fresh HMAC of that delivery's body. Because HMAC avalanches, two pings a quarter of an hour apart, differing only in sent_at, share no visible structure at all. A signature that never changed would be a replay token, not a signature. If the header ever looked like the secret, that would be the bug.
"TypeError: the key argument must be of type string"
Received undefined. This is not a crypto problem: the secret variable resolved to nothing before hashing began. On Pipedream, process.env.YOUR_VAR is empty until you create that variable in workspace settings, which is exactly why the snippet above takes the secret as a step prop instead. Guard the read and throw a message that names the missing variable, so the next failure is a sentence rather than a stack trace.
The transferable lesson: a signature mismatch is not one failure, it is four, and they are indistinguishable from the outside. Wrong key, wrong bytes, wrong comparison, missing config. A verifier that just returns false is a verifier you will debug by guessing. Print the expected value, the received value, and the body length on failure, and all four become one glance.
Section 08 · retries, rate limiting, and giving up
Your server sets the pace, and we stop when you stop answering.
A delivery engine that ignores backpressure is a denial-of-service tool pointed at your own customers. Three rules keep it polite: retry only what is worth retrying, wait as long as the receiver asks, and stop entirely when an endpoint has been dead for a while.
flowchart TB
P["POST attempt n"] --> C{"outcome"}
C -->|"2xx"| OK["success · streak reset to 0"]
C -->|"other 4xx"| PERM["permanent · refused, not retried"]
C -->|"408 · 429 · 5xx · transport error"| R{"attempt < 3?"}
R -->|"no"| EX["exhausted"]
R -->|"yes"| W{"Retry-After header?"}
W -->|"present, ≤ 30s"| WA["wait exactly that long"]
W -->|"present, > 30s"| EX
W -->|"absent"| WB["wait 2s, then 4s"]
WA --> P
WB --> P
PERM --> EX
EX --> S["consecutive_failures + 1"]
S --> D{"streak ≥ 20?"}
D -->|"no"| KEEP["stays active"]
D -->|"yes"| OFF["is_active = false
disabled_at · disabled_reason"]
What counts as worth retrying
Resending identical bytes can only help if the receiver's answer might change. A 500 might; a 401 will not. So the classification is small and deliberate, and the two codes a rate limiter actually speaks are both on the retryable side.
// internal/webhooks/service.go
// 429 and 503 are the two codes a rate limiter actually speaks, and both are
// retryable by definition: they mean "later", not "no". Every other 4xx is a
// refusal of these exact bytes, and a 401 will not start accepting an
// unchanged signature however many times we resend it.
func classify(code int) outcome {
switch {
case code == http.StatusRequestTimeout, // 408
code == http.StatusTooManyRequests: // 429
return outcomeRetryable
case code >= 500:
return outcomeRetryable
default:
return outcomePermanent
}
} Which rate-limit headers we respect
Exactly one: the standard Retry-After, in both forms RFC 9110 permits. When it is present and within the cap, it replaces our backoff entirely. That ordering is the whole point of reading it: a rate-limited endpoint knows when its window resets, and we are guessing.
// Both RFC 9110 forms are accepted: delta-seconds and an HTTP-date. Retry-After: 120 Retry-After: Wed, 21 Oct 2026 07:28:00 GMT // A past date means "you may retry now" and clamps to zero. // A negative or unparseable value is ignored, and the exponential // fallback (2s, then 4s) applies instead. // // X-RateLimit-Remaining / X-RateLimit-Reset are NOT read. They are a // convention, not a standard: Reset is a unix timestamp for some vendors, // a delta in seconds for others, and milliseconds for a few. Guessing // wrong turns a one-second pause into a one-day one.
Why there is a 30-second cap
A honoured Retry-After holds a goroutine open. A receiver asking for ten minutes is telling us it is not coming back inside any window worth blocking on, so we record the request, stop retrying that event, and let the failure streak carry the signal instead. The number it asked for is stored on the delivery row either way, so the log shows what your server said even when we could not wait that long.
Why the X-RateLimit family is ignored
Those headers are a convention, not a standard. X-RateLimit-Reset is a unix timestamp for GitHub, a delta in seconds for others, and milliseconds for a few more, with no way to tell from the value alone. Reading it wrong turns a one-second pause into a one-day one. If you want us to wait, send Retry-After; a 429 without one still gets the exponential fallback.
When we stop sending altogether
Twenty consecutive exhausted deliveries auto-disable the endpoint. The next event is never dispatched to it, because the dispatcher's lookup filters on is_active, so a dead URL stops costing anyone anything. The threshold is deliberately large: a webhook is somebody's integration, and killing it over a brief outage is worse than dropping a handful of events. At three attempts each, twenty is sixty failed POSTs.
| Signal | What it means |
|---|---|
consecutive_failures | Exhausted deliveries since the last success. Any successful delivery resets it to zero, so it measures a current outage rather than lifetime history. |
disabled_at + disabled_reason | Stamped only on the active to inactive transition made by the trip wire, and carrying the last error. Their presence is how you tell "the platform gave up" from "an operator switched it off". |
| Recovering | Fix the endpoint, use the ping to confirm it answers, then re-enable. Re-enabling clears the streak and the stamp, otherwise a webhook one failure from the threshold would trip again immediately and look like the fix never took. |
| The ping is exempt | A test ping makes one attempt and touches the streak in neither direction. Pinging a dead endpoint must not push it toward auto-disable, and pinging a healthy-looking one must not paper over real deliveries that are failing. |
The rest of the contract
| Property | Behaviour |
|---|---|
| Transport | HTTPS only. An http:// URL is rejected at registration, not at delivery. |
| Timeout | 5 seconds per attempt. Return 2xx fast and do the work asynchronously on your side. |
| Duplicates | Retries mean at-least-once, not exactly-once: a receiver that times out after doing the work still gets a second copy. The envelope is byte-identical across attempts, so its signature is a usable idempotency key. |
| Ordering | Not guaranteed, and retries widen the gap. Each subscription is dispatched on its own goroutine, so two events emitted in order can land out of order. Use sent_at if order matters. |
| Delivery log | One row per attempt, carrying the attempt number, status code, duration, error snippet, and any Retry-After the receiver sent. Newest first, 50 by default and 200 at most. |
| Error capture | Up to 512 bytes of a non-2xx response body is kept, enough to see "401 unauthorized" or a short JSON error without turning the log into a data store. |
| Replay protection | Not yet. There is no timestamp header, so the signature proves authenticity, not freshness, and a captured delivery stays valid. Dedupe on payload contents or reject a stale sent_at yourself if that matters to you. |
The one remaining "not yet" is deliberate. Replay protection needs the timestamp inside the signature rather than beside it, which changes what receivers hash and therefore breaks every verifier written against today's contract. That is a versioned envelope change, not a patch, and shipping it badly is worse than not having it.
Section 09 · implementation surface
Where it lives.
One Go package, two migrations, one console screen. The HTTP surface is mounted under the workspace resolver alongside automations, gated by the same role handlers.
Middleware (Go)
internal/webhooks/service.go: sign, the retry loop, classification,Retry-Afterparsing, the ping pathinternal/webhooks/registry.go: the closed event catalog, mirroring automationsinternal/webhooks/repo.go: CRUD, the delivery log, and the single-statement circuit breakercmd/server/wiring/webhooks.go: theEventSinkadapter onto the automations publisherdb/migrations/000044_webhooks.up.sql: both tables and their indexesdb/migrations/000124_webhook_delivery_health.up.sql: failure streak, disable stamp, per-attempt columns
HTTP surface and console
GET · POST /api/v1/workspace/:handle/webhooks: list, createGET · PATCH · DELETE .../webhooks/:id: read, patch, removeGET .../webhooks/:id/deliveries: recent attempts, newest firstPOST .../webhooks/:id/test: synchronous ping, recorded like any delivery- Console: the Webhooks screen mints the secret client-side, reveals it once, and surfaces the failure streak plus any auto-disable banner
Sources & further reading
Reads and writes are role-gated: any member can list webhooks and read the delivery log; creating, patching, deleting, and pinging require manager. Cross-workspace ids resolve to not-found rather than leaking another tenant's delivery history.
internal/webhooks/: the package doc states what the subsystem owns and why the catalog is redeclareddb/migrations/000044_webhooks.up.sql: the schema, with the plaintext-secret rationale in the header commentinternal/webhooks/service_test.go: secret generation, redaction, fan-out, and the engine suite (classification, bothRetry-Afterforms, the cap, the threshold, and the ping's exemption from the streak)
Related engineering reading: the Link Bouncer, which solves the mirror-image problem of HMAC verification across two languages, Background Tasks for the queue platform that retries are likely to ride on, and the Dashboard & Telemetry Engine.