Skip to main content
Omazy Engineering

Image RAG · vision chat + visual catalog search

Seeing the catalog.

A customer photographs a product and asks whether the store has it. Giving a chat agent that ability took three composable layers: attach the image to the turn so the vision model reads it, describe the image into the catalog retrieval query, and match the photo itself against every product image with CLIP. Each layer degrades cleanly into the one below, and the embedding vendor sits behind a swappable, operator-managed driver.

// gpt-4o vision · CLIP 768-d · catalog_items.image_vec · brute-force cosine // RFC 2220 · shipped on main · deployed · verified in prod
768-d
CLIP vector per product image
4,938
catalog images embedded (of 4,939)
~10 ms
brute-force cosine match, per agent
0.55
cosine floor for a visual match
3
layers of image understanding

Section 01 · the problem

A photo that went nowhere.

A visitor to a storefront agent uploaded a product photo, typed "do you have this item?", and got a confident answer about a completely unrelated product. Worse, on reload their photo was gone from the transcript. Two bugs, one root cause.

The widget send path was text-only. The uploaded image went through a separate presigned-upload flow that persisted it as a file, but the chat message it accompanied carried only text. So the image was never linked to the turn: it did not survive into history, and the assistant's TextContent.Images array was empty, which meant the vision pipeline, already wired end to end downstream, was being fed nothing.

That last part reframed the whole effort. The hard machinery, a vision-capable model route, base64 inlining, multimodal message parts, already existed. The first fix was plumbing, not AI. The interesting work was what came after: turning "the agent can see the image" into "the agent finds that product in the catalog".

Section 02 · system architecture

One message, two understandings, rejoined.

An image turn fans out from a single visitor message into a vision answer and a visual match, running in parallel, then rejoins as a product carousel. Product vectors are built offline by a backfill job, so the online path is a pure similarity lookup.

flowchart TB
  V["Visitor
uploads a product photo"] --> W["Widget · Preact
upload, finalize, send with file_id"] W --> WW["widgetweb
persist TextContent.Images"] WW --> D["OUchat dispatcher
resolveImages to base64"] subgraph ONLINE["online · one image turn"] VIS["gpt-4o vision
reads image, writes reply"] EMB["ImageEmbedder · CLIP
photo to 768-d vector"] COS["Visual search
brute-force cosine ≥ 0.55"] end subgraph OFFLINE["offline · backfill"] BF["Backfill job
each product image to CLIP"] end D --> VIS D --> EMB --> COS COS -->|read vectors| PG[("Postgres
catalog_items.image_vec")] VIS --> CAR["Product carousel
top matches to widget"] COS --> CAR CAR --> V BF -->|write vectors| PG
Fig 1. A visitor message reaches the dispatcher, which forks it: the vision model reads the image and writes the reply, while the CLIP embedder turns the photo into a vector that is cosine-ranked against the agent's product vectors. Both branches rejoin as one carousel. The offline backfill is the only writer of catalog_items.image_vec.

Section 03 · layer 01

Attach the image to the turn.

The fix that unblocked everything. We widened the visitor send to carry an images[] array and persisted it onto the message as TextContent.Images: each entry a file_id for the model plus a display url for the transcript. The frontend now uploads first, then sends the turn with the finalized id.

Because the dispatcher already reads TextContent.Images, resolves each file to bytes, base64-inlines them, and routes any image turn to a vision model, populating that one array lit up the whole path. The thumbnail renders on reload and the assistant genuinely sees the picture. No new model work: the array was simply never being filled.

Result: the image persists in history, and the assistant reads it via gpt-4o vision. Two reported bugs, fixed by one wiring change on a backward-compatible send.

Section 04 · layer 02

Describe the image into the retrieval query.

Seeing the image is not the same as finding the product. The catalog retriever keys off the user's text, and "do you have this item?" retrieves nothing useful.

So on an image turn we ask the vision model for a compact, search-oriented description (brand, type, colour, attributes) and fold it into a separate retrievalQuery that drives the existing hybrid catalog RAG. The answering model is untouched; only what we retrieve gets smarter. This is the cheapest useful layer: it reuses the whole text-embedding and BM25 retrieval stack and needs no new infrastructure. It is also the graceful fallback for the layer below. If visual similarity is unavailable or finds nothing, the vision-described query still surfaces relevant products.

Section 05 · layer 03

True visual similarity with CLIP.

The headline capability: match the photo itself to product images. We embed every catalog image and the uploaded photo into the same CLIP space, then rank by cosine similarity. The same product from a different angle lands as a near neighbour; unrelated items fall below a floor and are dropped.

sequenceDiagram
  autonumber
  participant Wg as Widget
  participant Ww as widgetweb
  participant Ds as Dispatcher
  participant Vs as gpt-4o vision
  participant Cl as CLIP · Replicate
  participant Pg as Postgres
  Wg->>Ww: send turn { text, images[file_id] }
  Ww->>Ds: persist message, trigger reply
  Ds->>Ds: resolveImages to base64 data URIs
  par Vision answer
    Ds->>Vs: prompt + image
    Vs-->>Ds: streamed reply
  and Visual match
    Ds->>Cl: EmbedImage(photo)
    Cl-->>Ds: 768-d vector
    Ds->>Pg: load agent image_vec rows
    Pg-->>Ds: vectors
    Ds->>Ds: cosine rank, keep ≥ 0.55
  end
  Ds-->>Wg: reply + product carousel
Fig 2. An image turn, end to end. The vision answer and the visual match run as parallel branches (par), then rejoin. Only AI turns carry an image, so a text turn never touches the CLIP or Postgres path.

Check the vendor before designing the index

The instinct was to reuse our Cloudflare Workers AI gateway. We checked the model catalogue before writing a line of index code, and Workers AI has no image-embedding model (image-to-text and classification, yes; a CLIP encoder, no). That single confirmation redrew the layer: it genuinely needs an external vendor. We landed on Replicate's krthr/clip-embeddings (CLIP ViT-L/14, 768-d), reached through its versioned predictions endpoint with a synchronous Prefer: wait.

Postgres and brute force, on purpose

The obvious move is a vector index, pgvector or OpenSearch kNN. We chose neither. Visual search is scoped per agent, and a storefront catalogue is a few thousand products. At that scale, storing each vector as JSONB on catalog_items.image_vec and computing cosine in Go over the agent's rows is a sub-10 ms loop with zero new infrastructure: no extension, no index to build and reindex, no mapping to migrate.

The whole ranking is a few lines:

// brute-force cosine over one agent's product vectors, fast at catalog scale
func cosineSim(a, b []float32) float64 {
    var dot, na, nb float64
    for i := range a {
        dot += float64(a[i]) * float64(b[i])
        na  += float64(a[i]) * float64(a[i])
        nb  += float64(b[i]) * float64(b[i])
    }
    if na == 0 || nb == 0 {
        return 0
    }
    return dot / (math.Sqrt(na) * math.Sqrt(nb))
}

The trade is honest: brute force is O(catalog) per query and will not hold at cross-tenant or six-figure-catalogue scale. But the whole thing sits behind an interface, so swapping in an ANN index later is a driver change, not a rewrite. Ship the simple thing that is correct today; keep the seam for tomorrow.

Section 06 · a swappable embedder

Managed by operators, not by an env file.

Every external model in the platform sits behind a narrow, name-keyed driver, and image embedding is no exception. ImageEmbedder is a three-method interface with a factory (IMAGE_EMBED_DRIVER) that defaults to disabled, so the feature is inert until a vendor is configured and never branches on a vendor name in the callers.

type ImageEmbedder interface {
    Name() string
    Dimensions() int
    EmbedImage(ctx context.Context, mimeType string, data []byte) ([]float32, error)
}
// factory: replicate | logged | disabled · callers depend only on the interface

Operator-managed config

A single-row llm_image_embed_config table holds the driver, model, dimensions, and an envelope-encrypted API key, stored exactly like every chat-provider credential and never echoed back. An operator sets it from the console; on first boot the running env value is seeded into the row so the live vendor is visible in the UI.

Hot reload

Consumers hold a LiveImageEmbedder, a holder that atomically swaps the underlying driver on save. Change the model or rotate the key in the UI and it applies live: no restart, no consumer rewiring. The Replicate driver caches its resolved model version, so a model change builds a fresh instance, which the swap does for free.

Coupled invariant: a cosine match requires the query vector and the stored vectors to share dimensions and model. Changing the model from the UI invalidates every stored image_vec; the config records image_vec_model precisely so a change can trigger a re-embed rather than silently returning garbage neighbours.

Section 07 · backfilling the corpus

Embedding 4,939 product images.

The online path is only as good as the corpus behind it, so we embedded every product image for the pilot agent. A resumable job pulls the catalogue, embeds each image through the driver, and writes the vector back, skipping rows already done so it can be re-run safely.

Incident and fix: at 6 concurrent Replicate predictions the failure rate hit about 58%, the account's concurrency ceiling surfacing as 429s. Dropping to 3 workers with exponential-backoff retry took failures to zero. A single serial probe had already proven the model worked; the lesson was purely about pacing a rate-limited vendor. Final coverage: 4,938 of 4,939 (one genuinely unfetchable image).

Section 08 · the dividend

Demand-gap tracking falls out for free.

Once agents could tell whether a wanted product was in the catalogue, a second capability came almost for free: recording the things customers ask for that we do not stock. We built it as a passive hook, not a conversational journey, because tracking should be invisible and never a slot-filling detour.

The signal differs by modality. For an uploaded photo, the visual-similarity floor makes an empty result a reliable "not stocked" marker. For text, an empty catalog search is not reliable, since semantic search returns nearest neighbours for anything, so a cheap async model judges the query against the products actually shown and, only if there is a genuine gap, names the missing product. Both feed a deduplicated catalog_gap_queries table surfaced as a "Demand gaps" report with request counts and CSV export. Merchandising now sees, in the customer's own words, what to stock next.

Section 09 · what shipped

Every layer is live.

ComponentWhereState
Image on the message + vision routewidgetweb, chats, widgetlive
Vision-describe → retrieval queryOUchat dispatcherlive
CLIP visual similarityReplicate · catalog_items.image_veclive
Modular ImageEmbedder + operator config + hot reloadinternal/llm, admin API, consolelive
Catalog vector backfill (4,938 images)resumable jobcomplete
Demand-gap tracking + reportdispatcher hook, catalog pagelive

Verified end to end in production: a photo of Phantom in Red Parfum Elixir returned that exact product as the top carousel card, alongside five visually related fragrances.

Section 10 · decisions and lessons

What we would tell the next team.

Confirm the vendor before designing the index

One catalogue lookup ruled out the obvious Workers AI reuse and redrew the phase. Verify the external capability exists before building on the assumption that it does.

Every capability degrades into the one below

CLIP match, then vision-described retrieval, then the model still sees and answers. No single dependency is load-bearing for a useful reply.

Backend first, backward-compatible

The widened send stayed compatible with the old API, so the backend could roll out before the new widget. No window where a new client hit an old server.

Async judge, not an agentic tool

A built-in tool would force the tool-calling loop on every turn and kill live token streaming. A detached classifier gives the same precision with zero visitor latency.

"Empty search" is not "not stocked"

Semantic retrieval returns neighbours for anything. For a reliable gap signal, use the visual floor for photos and an LLM judge for text.

Ship simple, keep the interface

Postgres plus brute-force cosine is correct at catalogue scale today; the ImageEmbedder and vector-store seams make an ANN upgrade a swap, not a rewrite.

Section 11 · implementation surface

Where it lives.

The pipeline spans the widget send, one dispatcher fork, a new driver package, and three small migrations. Nothing branches on a vendor name outside the factory.

Ingest + serve (online)

  • widget/, internal/widgetweb: image on the send, persisted to TextContent.Images
  • cmd/server/wiring/ouchat_dispatcher.go: resolveImages, vision route, retrieval-query fork
  • cmd/server/wiring/catalog_search.go: newVisualCatalogSearch + cosineSim
  • cmd/server/wiring/catalog_gaps.go: the passive demand-gap hook + async judge

Embedder + storage

  • internal/llm/image_embeddings.go: the ImageEmbedder interface, factory, LiveImageEmbedder
  • internal/llm/image_embeddings_replicate.go: the Replicate CLIP driver
  • internal/llm/image_embed_config.go: operator config, envelope-encrypted key
  • migrations 000112 / 000113 / 000114: image_vec, embed config, catalog_gap_queries

Sources & further reading

The three-layer plan and the vendor-confirmation decision are recorded in the RFC; the drivers carry their contracts in-file.

  • docs/rfc2220-widget-image-vision.md: the three layers, the vendor decision, the coupled dimension/model invariant
  • internal/llm/image_embeddings.go: the interface, the factory, the hot-swap holder
  • GitHub RFC 2216: the widget integration SDK that carries the image send end to end

Related engineering reading: the Omazy Harness (the runtime the dispatcher forks inside), and the Common Channel Wrapper (how the same message model reaches every channel the widget shares).