# Current Image Search page: 20-photo, eight-query pilot

The page uses the [pinned 6 October 2026 evaluation](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/3ca626712b086958963cfa5741b3993e1d81614e/evaluations/2026-10-06/image-search-0fe605cb2953dd92). It contains the exact prepared photos, original queries, output-blind relevance grades, caption recordings, embeddings, settings, complete ranks, usage and an offline verifier. Its companion records the corrected model-cost forecast.

The selected model is google/siglip-so400m-patch14-384, revision 9fdffc58afc957d1a03a25b10dba0329ab15c2a3, default profile. Every recipe received the same aspect-preserving JPEGs with a maximum 512-pixel long edge. Source listings, labels and gold were excluded from inputs. Eight purposive query/source families have 160 candidate relevance grades; the grades are not independent trials. Output-blind AI vision review supplied the grades. The page reports first-result counts, without a population parity or superiority claim.

For the same blue-shoe query, SigLIP 384 returns cat-01 first and GPT-6 Luna captions + text-embedding-3-small returns cat-03 first. That photo is visibly white. Luna's caption of cat-01 includes its blue colour and white sole; this case does not establish that captioning omitted the colour. Voyage 3.5 and GPT-6.1 Sol captions + text-embedding-3-large also matched all eight first results. The base-224 native variant matched seven and remains in the full evidence.

At the [published SIE Cloud list prices](https://superlinked.com/cloud), 100,000 image encodings cost $2.320733333334. The full forecast is one 100,000-photo index and 1,000,000 short queries, including fixed 64 padded native tokens per query: $4.301092445014 for SigLIP 384, $12.80736 for Voyage 3.5, $7.08 for Luna + small embeddings, and $161.55135 for Sol 6.1 + large embeddings. Captions are generated once per indexed photo. The list-price model/API forecast uses the recorded size/token distribution, excluding shared vector-store, storage and application costs equally. There are no index refreshes or Batch discounts. Native admission-tokenizer lengths are not billing units.

Encode returns dense vectors. The example preserves the JPEG bytes, indexes in two batches of 10, then encodes the query separately with is_query=True. Cosine ranking runs in the caller; scores are similarities, not probabilities. A model change requires re-encoding the index. Copy and signup Run preserve the full 20-image candidate set. The static code excerpt omits repetitive filenames and the cosine helper while keeping native calls visible.

The consequence clipping cites [Baymard Institute's Ecommerce Search UX 2026](https://baymard.com/research-articles/ecommerce-search-query-types), updated 29 April 2026 and originally published 12 September 2024. Its research describes search friction and users treating poor results as the available selection. It does not compare these models or measure a SIE conversion uplift. The page quotes seven words and uses an authored consequence heading.

Photo permissions and source revisions remain per-file. See [pilot attribution](/reference/image-search/pilot/ATTRIBUTION.md) and [pilot licence metadata](/reference/image-search/pilot/licences.json). Prepared JPEGs have no embedded source metadata.

# Archived catalogue study

The material below records an earlier 309-query/2,573-photo study, retained for existing data consumers. It does not support the current page's pilot counts or current prices.

# Image search page sources

The /image-search page reports one pre-registered study, run on 30 September
2026: SIE's SigLIP so400m (`google/siglip-so400m-patch14-384`, live on SIE
Cloud) against the
hosted image search a shop already runs, on a held-out product catalogue and on
two public text-to-image sets. The design, the decision rules and the catalogue
were fixed before any tested model embedded a photo or a question, and amended
once, at the freeze, before any tested model ran. Every figure on the page is
regenerated from the recorded run by
`apps/site/scripts/import-image-search-evidence.mjs` and checked at build time
in `apps/site/src/data/reference/tasks/image-search.ts`.

## The catalogue

**Source.** Amazon Berkeley Objects (ABO), Amazon's public product catalogue
of 147,702 listings with their photos, licensed
[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)
([dataset](https://amazon-berkeley-objects.s3.amazonaws.com/index.html)).
The page shows ABO photos under that licence and credits them under the hero.

**Labels.** A listing entered the catalogue when ABO's own metadata gives it one
colour, one material and one product type, and its main photo is at least 500
pixels on its long side:

- ABO's standardised colour, folded to 15 words a shopper types. Multicoloured,
  clear and metallic listings were dropped.
- ABO's free-text material, folded by keyword to 13 families. Cotton,
  polyester, linen, wool and canvas are all "fabric", which is what they look
  like in a photo; faux leather is leather. A listing matching two families or
  none was dropped.
- ABO's product type, mapped to one noun; generic types and jewellery dropped.

That gave 4,571 listings.

**Blind label check.** Claude Sonnet 5 saw each photo at 512 pixels, without
its label, and named its colour, material and type from the same words. A photo
stayed only if all three answers matched its label. The checker is an
Anthropic model because no tested product is: a GPT checker would have shaped a
catalogue that a GPT captioner is then tested on. Words a photo cannot tell
apart were folded on both sides first (running shoes are shoes, a tote is a
handbag, a bed frame is a bed, mesh is fabric, a silicone case reads as
plastic). **2,573 photos** passed. 392 of them, drawn at random, were then
inspected by eye on labelled contact sheets; none was dropped, and three are
borderline on colour (two clocks and a rug).

**Questions.** A question is a label written as a shopper types it: "brown
leather sofa". It is templated, not written by any model. A photo is a right
answer when its checked label equals the question's: a shopper who types
"brown leather sofa" is served by any brown leather sofa. A label became a
question only if the catalogue also holds a near miss on at least two of its
three words (another photo sharing two words and differing on the third), so
first place has to be won over a near miss. That gave **309 questions**, a
number fixed by the catalogue. The catalogue and the questions were hashed and
the hashes committed before any tested model ran.

## The products

| Product | Called as | Photo | Query |
|---|---|---|---|
| SIE SigLIP so400m | `google/siglip-so400m-patch14-384` on the open-source SIE server, the checkpoint SIE Cloud serves. On the 60-question pilot SIE Cloud and the open-source server gave identical first results; the full run's figures are the open-source server's, because SIE Cloud's encode route was failing when the full question set ran | JPEG bytes | the question |
| SIE SigLIP so400m-224 | `google/siglip-so400m-patch14-224` on the open-source server; not on SIE Cloud yet | JPEG bytes | the question |
| Voyage multimodal-3.5 | Voyage `multimodalembeddings` | `input_type: "document"` | `input_type: "query"` |
| Cohere Embed v4 | Cohere `v2/embed`, `embed-v4.0`, float | `search_document` | `search_query` |
| GPT-6 Luna captions + text-embedding-3-small | GPT-6 Luna captions each photo once at `detail: "low"`, then `text-embedding-3-small` embeds the caption | the caption | `text-embedding-3-small` |

OpenAI has no image embedding model, so an OpenAI shop captions each photo
with a GPT model and embeds the caption. The caption prompt gives that pipeline
its best case by asking for what a shopper names: "Describe this product in one
sentence for a shop's search index: say what it is, its main colour and its
main material."

Cohere Embed v5 is not in the study: Cohere publishes no per-unit price for it,
only hourly Model Vault deployments, so it cannot sit on a per-photo price
axis. Google Gemini Embedding 2, Vertex `multimodalembedding@001` and Amazon
Nova 2 Multimodal Embeddings were not run: the study had no credential for them.
The page makes no claim about them.

Every product ranks every photo by cosine similarity, photos at 1,024 pixels on
the long side.

## Results: the held-out catalogue

The figure is the share of questions whose first result has the colour, the
material and the product type the question names.

| Product | Right product first | % |
|---|---|---|
| Cohere Embed v4 | 260 of 309 | 84.1% |
| SIE SigLIP so400m-224 | 260 of 309 | 84.1% |
| SIE SigLIP so400m | 257 of 309 | 83.2% |
| Voyage multimodal-3.5 | 232 of 309 | 75.1% |
| GPT-6 Luna captions + text-embedding-3-small | 165 of 309 | 53.4% |

Pairwise, exact McNemar, with a paired bootstrap 95% interval on SIE minus the
other (10,000 draws):

- **Voyage multimodal-3.5:** 43 questions only SIE got right, 18 only Voyage
  did; SIE ahead by 8.1 points, interval 3.2 to 12.9, p = 0.002. This is the
  pre-registered bar for "more often than Voyage" (a lower bound above zero).
- **Cohere Embed v4:** 23 and 26; Cohere a point higher, interval -5.2 to 3.6,
  p = 0.78. That misses the pre-registered bar for "as often as" (a lower bound
  of at least -4), so the page makes no quality claim against Cohere either way.
  Its claim against Cohere is price.
- **GPT-6 Luna captions:** 106 and 14; SIE ahead by 29.8 points, p < 0.001.
- **SIE SigLIP so400m-224:** 9 and 6 against the 384, p = 0.61: level, at 2.5
  times the throughput (below).

The page illustrates **q159, "gold metal ceiling light"**, using both models'
actual first results from the same frozen query. Voyage multimodal-3.5 ranks
the black metal light (`91vfAt-ORNL`) first. SIE SigLIP so400m ranks the gold
metal light (`71HwFTkKWwL`) first. The requested colour is highlighted in the
query and the wrong colour on the Voyage result. The SIE result is unmarked.
The importer validates both image IDs against their recorded ranking files,
the query against `questions.json`, and the labels against `catalogue.json`.

This illustration was selected from the saved colour misses for its clear
photographic contrast and existing licensed display assets. It does not change
the held-out score or the pre-registered rule. All saved candidates remain in
the evidence packet, including:

- "white wood dresser": Voyage's first photo is a brown wood dresser; its first
  white one is second.
- "gold metal ceiling light": Voyage's first photo is a black one; its first
  gold one is third.
- "grey wood cabinet": the caption pipeline's first photo is a white wood
  cabinet; its first grey one is third.

## Results: rival resolution

Voyage and Cohere also embedded the catalogue at 512 and 224 pixels. Both are
best at 1,024: Voyage 73.1% at 512 and 70.2% at 224; Cohere 81.2% at 512 and
82.2% at 224. The chart shows each at its best score and prices each at its
cheapest photo, so neither the resolution nor the price flatters SIE.

## Results: public text-to-image sets

Flickr30k (1,000 images, 5,000 captions) and MS-COCO (5,000 images, 25,010
captions), the Karpathy test splits. Each caption is a query; its own image is
the one right answer. Right image first:

| Product | Flickr30k | MS-COCO | Mean |
|---|---|---|---|
| Cohere Embed v4 | 88.1% | 63.2% | 75.6% |
| Voyage multimodal-3.5 | 82.7% | 57.3% | 70.0% |
| SIE SigLIP so400m | 83.2% | 54.1% | 68.7% |
| SIE SigLIP so400m-224 | 75.4% | 51.9% | 63.6% |
| GPT-6 Luna captions + text-embedding-3-small | 22.9% | 12.0% | 17.4% |

On these scene photos Cohere and Voyage rank the right image first more often
than SIE (1.3 points behind Voyage and 7.0 behind Cohere on the mean). The pre-registration says the held-out catalogue
decides where the two disagree: it is the reader's own kind of photo, and
public sets may be in any model's training data. The page's claim is about
product photos and nothing wider. The caption arm's low score here is partly its
prompt, which asks about a product's colour and material and fits a street
scene badly.

## Price

The page prices one job: **index a million product photos**, each product at its
cheapest real-time photo. SIE has no batch tier, so the page compares real-time
with real-time. Rounding goes against SIE: its figure is rounded up, rivals'
down, and the page's "less per photo" is checked on the rounded figures.

| Product | Cheapest real-time photo | A million photos |
|---|---|---|
| SIE SigLIP so400m | $0.0232 per 1,000 images, SIE's live rate book | $24 |
| Cohere Embed v4 | 55.1 billed image tokens for a 224-pixel photo, at $0.47 per 1M image tokens | $25 |
| Voyage multimodal-3.5 | the 50,000-pixel floor, $0.60 per billion pixels: $0.00003 an image | $30 |
| GPT-6 Luna captions + text-embedding-3-small | the recorded caption tokens at $0.10 / $0.50 per 1M, plus the caption at 3-small's $0.02 per 1M | $37 |

Prices read on 30 September 2026:
[Voyage](https://docs.voyageai.com/docs/pricing),
[OpenAI](https://developers.openai.com/api/docs/pricing); Cohere's image-token
rate is the resold rate (cohere.com/pricing lists none), read on two resellers.
On their batch APIs Voyage would cost $20 and the caption pipeline $18 for the
same job. Voyage also gives each new account 150 billion free pixels, about 3
million photos at its floor, once; the job above is priced without it.

**SIE's price and cost.** On the open-source SIE server on one L4 GPU with 8 CPU
cores and 32 GiB, 1,024-pixel JPEGs at 8 photos a request and 16 requests in
flight, the 384 sustained 43.1 photos a second and the 224 sustained 110. At
Modal's list prices for that whole lane, the 1.75 region factor and 75%
utilisation, that is about $0.022 per 1,000 photos for the 384 and $0.0084 for
the 224. The 224 scores level with the 384 on the catalogue, and is queued for
SIE's rate book at $0.017 per 1,000, twice its cost; the page will move to it
when that price is live.

## A stronger open model (screened, not adopted)

Each candidate ran on the held-out catalogue. The pre-registered bar to replace
SigLIP was 3 points above the 384 with p < 0.05:

| Model | Right product first | Against so400m-384 |
|---|---|---|
| Meta PE-Core bigG-14-448 (through OpenCLIP; SIE cannot serve it yet) | 87.4% | +4.2 points, p = 0.06 |
| SigLIP 2 giant (opt, patch16-384) | 85.8% | +2.6, p = 0.24 |
| SigLIP 2 so400m-512 | 85.8% | +2.6, p = 0.23 |
| SigLIP 2 so400m-384 | 85.1% | +1.9, p = 0.38 |
| Qwen3-VL-Embedding-8B | 79.3% | -3.9, p = 0.17 |

None passed; PE-Core bigG came closest. It is the model to revisit once SIE can serve it (it needs
newer timm and OpenCLIP than SIE pins). Two screened models scored above 85%,
the level the pre-registration set for a discriminating test; the four products
the page compares all score under it.

## Latency

Not on the page. The pre-registered rule puts latency on the page only if
SIE's is the lowest. SIE Cloud's encode route was failing on 30 September, so
SIE's hosted latency was not measured; Voyage, Cohere and OpenAI answered a
text query in 130 to 260 ms.

## The runnable example

[`examples/image-search`](https://github.com/superlinked/sie/tree/395366cd58d0c14eddfd59c84daec438acb5d8c8/examples/image-search)
fetches the frozen catalogue, the questions, SIE's recorded vectors and every
product's recorded rankings from
[superlinked/sie-task-evidence at revision `96c541a`](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/96c541a6a72ecb2ca72ee3644caecdb209ed476c/image-search)
and reproduces every figure above with no key.
