Image RAG · vision chat + visual catalog search
Seeing the catalog.
A customer photographs a product and asks whether the store has it. Giving a chat agent that ability took three composable layers: attach the image to the turn so the vision model reads it, describe the image into the catalog retrieval query, and match the photo itself against every product image with CLIP. Each layer degrades cleanly into the one below, and the embedding vendor sits behind a swappable, operator-managed driver.
Section 01 · the problem
A photo that went nowhere.
A visitor to a storefront agent uploaded a product photo, typed "do you have this item?", and got a confident answer about a completely unrelated product. Worse, on reload their photo was gone from the transcript. Two bugs, one root cause.
The widget send path was text-only. The uploaded image went through a separate presigned-upload flow that persisted it as a file, but the chat message it accompanied carried only text. So the image was never linked to the turn: it did not survive into history, and the assistant's TextContent.Images array was empty, which meant the vision pipeline, already wired end to end downstream, was being fed nothing.
That last part reframed the whole effort. The hard machinery, a vision-capable model route, base64 inlining, multimodal message parts, already existed. The first fix was plumbing, not AI. The interesting work was what came after: turning "the agent can see the image" into "the agent finds that product in the catalog".
Section 02 · system architecture
One message, two understandings, rejoined.
An image turn fans out from a single visitor message into a vision answer and a visual match, running in parallel, then rejoins as a product carousel. Product vectors are built offline by a backfill job, so the online path is a pure similarity lookup.
flowchart TB V["Visitor
uploads a product photo"] --> W["Widget · Preact
upload, finalize, send with file_id"] W --> WW["widgetweb
persist TextContent.Images"] WW --> D["OUchat dispatcher
resolveImages to base64"] subgraph ONLINE["online · one image turn"] VIS["gpt-4o vision
reads image, writes reply"] EMB["ImageEmbedder · CLIP
photo to 768-d vector"] COS["Visual search
brute-force cosine ≥ 0.55"] end subgraph OFFLINE["offline · backfill"] BF["Backfill job
each product image to CLIP"] end D --> VIS D --> EMB --> COS COS -->|read vectors| PG[("Postgres
catalog_items.image_vec")] VIS --> CAR["Product carousel
top matches to widget"] COS --> CAR CAR --> V BF -->|write vectors| PG
catalog_items.image_vec.Section 03 · layer 01
Attach the image to the turn.
The fix that unblocked everything. We widened the visitor send to carry an images[] array and persisted it onto the message as TextContent.Images: each entry a file_id for the model plus a display url for the transcript. The frontend now uploads first, then sends the turn with the finalized id.
Because the dispatcher already reads TextContent.Images, resolves each file to bytes, base64-inlines them, and routes any image turn to a vision model, populating that one array lit up the whole path. The thumbnail renders on reload and the assistant genuinely sees the picture. No new model work: the array was simply never being filled.
Result: the image persists in history, and the assistant reads it via gpt-4o vision. Two reported bugs, fixed by one wiring change on a backward-compatible send.
Section 04 · layer 02
Describe the image into the retrieval query.
Seeing the image is not the same as finding the product. The catalog retriever keys off the user's text, and "do you have this item?" retrieves nothing useful.
So on an image turn we ask the vision model for a compact, search-oriented description (brand, type, colour, attributes) and fold it into a separate retrievalQuery that drives the existing hybrid catalog RAG. The answering model is untouched; only what we retrieve gets smarter. This is the cheapest useful layer: it reuses the whole text-embedding and BM25 retrieval stack and needs no new infrastructure. It is also the graceful fallback for the layer below. If visual similarity is unavailable or finds nothing, the vision-described query still surfaces relevant products.
Section 05 · layer 03
True visual similarity with CLIP.
The headline capability: match the photo itself to product images. We embed every catalog image and the uploaded photo into the same CLIP space, then rank by cosine similarity. The same product from a different angle lands as a near neighbour; unrelated items fall below a floor and are dropped.
sequenceDiagram
autonumber
participant Wg as Widget
participant Ww as widgetweb
participant Ds as Dispatcher
participant Vs as gpt-4o vision
participant Cl as CLIP · Replicate
participant Pg as Postgres
Wg->>Ww: send turn { text, images[file_id] }
Ww->>Ds: persist message, trigger reply
Ds->>Ds: resolveImages to base64 data URIs
par Vision answer
Ds->>Vs: prompt + image
Vs-->>Ds: streamed reply
and Visual match
Ds->>Cl: EmbedImage(photo)
Cl-->>Ds: 768-d vector
Ds->>Pg: load agent image_vec rows
Pg-->>Ds: vectors
Ds->>Ds: cosine rank, keep ≥ 0.55
end
Ds-->>Wg: reply + product carousel
par), then rejoin. Only AI turns carry an image, so a text turn never touches the CLIP or Postgres path.Check the vendor before designing the index
The instinct was to reuse our Cloudflare Workers AI gateway. We checked the model catalogue before writing a line of index code, and Workers AI has no image-embedding model (image-to-text and classification, yes; a CLIP encoder, no). That single confirmation redrew the layer: it genuinely needs an external vendor. We landed on Replicate's krthr/clip-embeddings (CLIP ViT-L/14, 768-d), reached through its versioned predictions endpoint with a synchronous Prefer: wait.
Postgres and brute force, on purpose
The obvious move is a vector index, pgvector or OpenSearch kNN. We chose neither. Visual search is scoped per agent, and a storefront catalogue is a few thousand products. At that scale, storing each vector as JSONB on catalog_items.image_vec and computing cosine in Go over the agent's rows is a sub-10 ms loop with zero new infrastructure: no extension, no index to build and reindex, no mapping to migrate.
The whole ranking is a few lines:
// brute-force cosine over one agent's product vectors, fast at catalog scale
func cosineSim(a, b []float32) float64 {
var dot, na, nb float64
for i := range a {
dot += float64(a[i]) * float64(b[i])
na += float64(a[i]) * float64(a[i])
nb += float64(b[i]) * float64(b[i])
}
if na == 0 || nb == 0 {
return 0
}
return dot / (math.Sqrt(na) * math.Sqrt(nb))
} The trade is honest: brute force is O(catalog) per query and will not hold at cross-tenant or six-figure-catalogue scale. But the whole thing sits behind an interface, so swapping in an ANN index later is a driver change, not a rewrite. Ship the simple thing that is correct today; keep the seam for tomorrow.
Section 06 · a swappable embedder
Managed by operators, not by an env file.
Every external model in the platform sits behind a narrow, name-keyed driver, and image embedding is no exception. ImageEmbedder is a three-method interface with a factory (IMAGE_EMBED_DRIVER) that defaults to disabled, so the feature is inert until a vendor is configured and never branches on a vendor name in the callers.
type ImageEmbedder interface {
Name() string
Dimensions() int
EmbedImage(ctx context.Context, mimeType string, data []byte) ([]float32, error)
}
// factory: replicate | logged | disabled · callers depend only on the interface Operator-managed config
A single-row llm_image_embed_config table holds the driver, model, dimensions, and an envelope-encrypted API key, stored exactly like every chat-provider credential and never echoed back. An operator sets it from the console; on first boot the running env value is seeded into the row so the live vendor is visible in the UI.
Hot reload
Consumers hold a LiveImageEmbedder, a holder that atomically swaps the underlying driver on save. Change the model or rotate the key in the UI and it applies live: no restart, no consumer rewiring. The Replicate driver caches its resolved model version, so a model change builds a fresh instance, which the swap does for free.
Coupled invariant: a cosine match requires the query vector and the stored vectors to share dimensions and model. Changing the model from the UI invalidates every stored image_vec; the config records image_vec_model precisely so a change can trigger a re-embed rather than silently returning garbage neighbours.
Section 07 · backfilling the corpus
Embedding 4,939 product images.
The online path is only as good as the corpus behind it, so we embedded every product image for the pilot agent. A resumable job pulls the catalogue, embeds each image through the driver, and writes the vector back, skipping rows already done so it can be re-run safely.
Incident and fix: at 6 concurrent Replicate predictions the failure rate hit about 58%, the account's concurrency ceiling surfacing as 429s. Dropping to 3 workers with exponential-backoff retry took failures to zero. A single serial probe had already proven the model worked; the lesson was purely about pacing a rate-limited vendor. Final coverage: 4,938 of 4,939 (one genuinely unfetchable image).
Section 08 · the dividend
Demand-gap tracking falls out for free.
Once agents could tell whether a wanted product was in the catalogue, a second capability came almost for free: recording the things customers ask for that we do not stock. We built it as a passive hook, not a conversational journey, because tracking should be invisible and never a slot-filling detour.
The signal differs by modality. For an uploaded photo, the visual-similarity floor makes an empty result a reliable "not stocked" marker. For text, an empty catalog search is not reliable, since semantic search returns nearest neighbours for anything, so a cheap async model judges the query against the products actually shown and, only if there is a genuine gap, names the missing product. Both feed a deduplicated catalog_gap_queries table surfaced as a "Demand gaps" report with request counts and CSV export. Merchandising now sees, in the customer's own words, what to stock next.
Section 09 · what shipped
Every layer is live.
| Component | Where | State |
|---|---|---|
| Image on the message + vision route | widgetweb, chats, widget | live |
| Vision-describe → retrieval query | OUchat dispatcher | live |
| CLIP visual similarity | Replicate · catalog_items.image_vec | live |
Modular ImageEmbedder + operator config + hot reload | internal/llm, admin API, console | live |
| Catalog vector backfill (4,938 images) | resumable job | complete |
| Demand-gap tracking + report | dispatcher hook, catalog page | live |
Verified end to end in production: a photo of Phantom in Red Parfum Elixir returned that exact product as the top carousel card, alongside five visually related fragrances.
Section 10 · decisions and lessons
What we would tell the next team.
Confirm the vendor before designing the index
One catalogue lookup ruled out the obvious Workers AI reuse and redrew the phase. Verify the external capability exists before building on the assumption that it does.
Every capability degrades into the one below
CLIP match, then vision-described retrieval, then the model still sees and answers. No single dependency is load-bearing for a useful reply.
Backend first, backward-compatible
The widened send stayed compatible with the old API, so the backend could roll out before the new widget. No window where a new client hit an old server.
Async judge, not an agentic tool
A built-in tool would force the tool-calling loop on every turn and kill live token streaming. A detached classifier gives the same precision with zero visitor latency.
"Empty search" is not "not stocked"
Semantic retrieval returns neighbours for anything. For a reliable gap signal, use the visual floor for photos and an LLM judge for text.
Ship simple, keep the interface
Postgres plus brute-force cosine is correct at catalogue scale today; the ImageEmbedder and vector-store seams make an ANN upgrade a swap, not a rewrite.
Section 11 · implementation surface
Where it lives.
The pipeline spans the widget send, one dispatcher fork, a new driver package, and three small migrations. Nothing branches on a vendor name outside the factory.
Ingest + serve (online)
widget/,internal/widgetweb: image on the send, persisted toTextContent.Imagescmd/server/wiring/ouchat_dispatcher.go: resolveImages, vision route, retrieval-query forkcmd/server/wiring/catalog_search.go:newVisualCatalogSearch+cosineSimcmd/server/wiring/catalog_gaps.go: the passive demand-gap hook + async judge
Embedder + storage
internal/llm/image_embeddings.go: theImageEmbedderinterface, factory,LiveImageEmbedderinternal/llm/image_embeddings_replicate.go: the Replicate CLIP driverinternal/llm/image_embed_config.go: operator config, envelope-encrypted keymigrations 000112 / 000113 / 000114:image_vec, embed config,catalog_gap_queries
Sources & further reading
The three-layer plan and the vendor-confirmation decision are recorded in the RFC; the drivers carry their contracts in-file.
docs/rfc2220-widget-image-vision.md: the three layers, the vendor decision, the coupled dimension/model invariantinternal/llm/image_embeddings.go: the interface, the factory, the hot-swap holder- GitHub RFC 2216: the widget integration SDK that carries the image send end to end
Related engineering reading: the Omazy Harness (the runtime the dispatcher forks inside), and the Common Channel Wrapper (how the same message model reaches every channel the widget shares).