# Visual document search evidence

This file backs every figure on [/visual-document-search](https://superlinked.com/visual-document-search). It has
three parts. Only the first two describe what the page shows.

1. **The current list-price claim.** The 80% in the headline, the proof heading and the meta description is derived
   from published list prices applied to one recorded workload. The build recomputes it from the rate book.
2. **Frozen recorded runs the page shows.** The workload's usage, the retrieval-quality chart, the hero, the API example
   and the pain example come from records dated 5 September to 9 October 2026. Their numbers do not change.
3. **Frozen studies and earlier page editions.** The ViDoRe v3 study of 30 September 2026 and an earlier edition's
   closing diagram stay here for the record. The page shows none of their figures.

The [native-run notes](foundations/SOURCES.md) give more detail on parts 1 and 2.

## 1. The current list-price claim

### What the page says

- Headline: "Embed PDF pages for 80% less than Voyage at list prices"
- Subtitle and meta description: "Search tables and charts from page images with open models. At list prices, our
  256-page, 24-query workload costs 80% less on SIE than on Voyage 3.5."
- Proof heading: "Index and query PDF pages for 80% less than Voyage at list prices", above the two embedding API
  bills: SIE $0.053 and Voyage 3.5 images $0.286.

All three state one figure. At published list prices, the recorded workload below costs 81.5063% less to embed on SIE
than on Voyage multimodal-3.5. The page rounds that down to 80%.

### What it is, and what it is not

It is a list-price comparison: recorded usage multiplied by published per-unit prices, the same way for both
providers. It is not an invoice, a measured hosted bill or a measured native saving. The native run's
[`summary.json`](https://huggingface.co/datasets/superlinked/sie-task-evidence/blob/f770419eeadf46068ca3894816fd533f217371d1/visual-document-search/2026-10-09/native-and-saved-providers/summary.json)
records `material_advantage_established: false`, and leaves a complete measured saving unresolved, because none of
SIE's 33 image-index replies carries a usage record. The comparison does not fill that gap with an estimate. It prices
SIE's index calls by image count, the unit the price list bills for ordinary image-only encoding.

The 80% says nothing about retrieval quality. Quality on the same workload is in part 2, and no winner is established.

### Inputs

**Workload (frozen).** The primary frame of the 6 October 2026 provider screen and the 9 October 2026 native run:
256 candidate page images from 24 independent publication families, and 24 text questions, one per family. The four
correlated IRS questions and their eight pages are excluded on both sides.

- SIE, from the native run's
  [`native-timing-and-usage.json`](https://huggingface.co/datasets/superlinked/sie-task-evidence/blob/f770419eeadf46068ca3894816fd533f217371d1/visual-document-search/2026-10-09/native-and-saved-providers/native-timing-and-usage.json)
  (`primary256_pages24_queries`): 32 index calls of 8 page images, and 24 query calls that processed 832 input tokens
  in total, from 20 to 49 a call.
- Voyage multimodal-3.5, from the provider screen's
  [`timing-and-usage.json`](https://huggingface.co/datasets/superlinked/sie-task-evidence/blob/252d73593a7c2f5dc90e72b254d48a2a66df6c0c/evaluations/2026-10-06/visual-document-search-provider-screen-results-v1/timing-and-usage.json)
  (primary cohort): 477,166,132 image pixels and 562 text tokens, as Voyage's own replies reported them.

**SIE list prices.** Rate book `2026-10-10-production-bootstrap-v28`, effective `2026-10-10T18:00:00Z`, rate-book
SHA-256 `ab292600c88c4d95dfb67f7f568f9845eb738752208c9c4ff3bbd64ef3e9455d`. The site publishes this book at
[/pricing/sie-cloud-prices.json](/pricing/sie-cloud-prices.json) and vendors it as `apps/site/src/data/rate-book.json`.

| SKU | Exact USD per unit | Per 1,000 or per million |
|---|---|---|
| `encode:TomoroAI/tomoro-colqwen3-embed-4b:compact:images:us` | `1 / 5000` per page image | $0.20 per 1,000 page images |
| `encode:TomoroAI/tomoro-colqwen3-embed-4b:compact:input_tokens:us` | `1 / 500000` per input token | $2.00 per million input tokens |

The dated quote certificate, `apps/site/src/data/reference/tasks/visual-document-search-current-price-evidence.json`
(9 October 2026), is historical evidence recorded against an earlier book, `2026-10-08-production-bootstrap-v23` (SHA-256
`dad3a71d574819a2b2a6106db545d7143512e3bbc1d6762b854f055ac1fbbb6c`). Both Tomoro compact rows are identical in v23 and
every later book through v28, so no figure changed when the site moved to v25, v26, v27 or v28. The page computes from v28.

**Voyage list prices.** `voyage-multimodal-3.5` at its real-time list price on
[docs.voyageai.com/docs/pricing](https://docs.voyageai.com/docs/pricing): $0.60 per billion image pixels and $0.12 per
million text tokens. The provider screen recorded these prices on 6 October 2026
([`tariffs.json`](https://huggingface.co/datasets/superlinked/sie-task-evidence/blob/252d73593a7c2f5dc90e72b254d48a2a66df6c0c/evaluations/2026-10-06/visual-document-search-provider-screen-results-v1/tariffs.json)),
and they were checked again on 9 October 2026. Free credits and account discounts are excluded on both sides.

### Computation

SIE bills in credits: $50 buys 5,000,000, so one credit is $0.00001. Each call's charge rounds up to a whole credit,
once per call. A page image is 20 credits; an input token is 0.2 credits.

| Stage | SIE | Voyage multimodal-3.5 |
|---|---|---|
| Index 256 page images | 32 calls × 160 credits = 5,120 credits = $0.05120 | 477,166,132 pixels × $0.60 / 10⁹ = $0.2862996792 |
| Encode 24 questions | 176 credits = $0.00176 (each call rounded up; the 832 tokens pooled would be 166.4) | 562 tokens × $0.12 / 10⁶ = $0.00006744 |
| Total | 5,296 credits = **$0.05296** (the chart shows $0.053) | **$0.2863671192** (the chart shows $0.286) |

Difference: 1 − 0.05296 / 0.2863671192 = 81.5063%.

**Rounding.** The headline takes the whole-number part of that percentage, 81, and rounds it down to a multiple of
5: 80%. SIE's per-call credits round up. Both round against SIE, and so do the chart's three-decimal bills: $0.05296
shows as $0.053, and $0.2863671192 as $0.286.

### What the bills exclude

Both bills cover the embedding API calls only. They exclude page rendering, ingestion, vector storage, caller-side
MaxSim ranking and any parsing or OCR. The run's optional top-five visual reranking and its supplied-text OpenAI
control stay in the evidence. They are different stages with different inputs, so the comparison leaves them out.

### Managed rates on the page

The page shows managed SIE Cloud rates in two places, both from the v28 rows above: the SIE bill in the proof chart,
and the Cloud card's 250,000 page images per $50 ($50 at $0.20 per 1,000 page images). The card's figure counts
indexing only; text queries are billed separately at the input-token rate.

### What the build checks

`apps/site/src/data/reference/tasks/visual-document-search-costs.ts` recomputes the SIE bill from the vendored rate
book on every build. The build fails if either Tomoro compact row is missing, if the credits or the dollar total differ
from the dated certificate, if the rounded difference falls below 40%, or if the native run missed its useful-quality
floor.

## 2. Frozen recorded runs the page shows

### Retrieval quality on the same workload

The proof chart's quality pane plots nDCG@10 over the 24 families. SIE's row comes from the
[9 October 2026 native and saved-provider run](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/f770419eeadf46068ca3894816fd533f217371d1/visual-document-search/2026-10-09/native-and-saved-providers)
(archive SHA-256 `82e89e2138a9c97331f20e8b0e2b44da3034b9b5e751e9bf5c18d1881131608e`). Voyage's row is the saved
6 October 2026 provider-screen result that the native run carries forward unchanged.

| Arm | nDCG@10 | Questions with a grade-2 page in the first five |
|---|---:|---:|
| Voyage multimodal-3.5 | 0.866 | 21 of 24 |
| SIE `TomoroAI/tomoro-colqwen3-embed-4b:compact` | 0.852 | 21 of 24 |

SIE minus Voyage is −0.013 nDCG@10. Its exploratory paired 95% family-bootstrap interval runs from −0.063 to +0.042
(10,000 replicates, seed 20261006). Every paired interval in the run crosses zero. The run establishes neither a
quality winner nor parity, and the page claims neither. It does show that SIE's compact model cleared the run's
pre-set useful floor: nDCG@10 of at least 0.65, and a grade-2 page in the first five for at least 20 of 24 questions.
Cohere and Gemini were unavailable; they are not plotted, and not counted as zero.

### Hero and API example: current-edition source-page illustration

The fresh illustration is published separately in [recorded 2026 task illustrations](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/74f4e0eb2676a95b625783932796edac12040644/task-visual-examples-2026). Its manifest SHA-256 is `662f2bbe970135bdc908556464de317a9e67be7a8398175bd11a8fbbd6de513b`. The generator verifies that pinned manifest and every file before deriving display data. The original aggregate study and runnable-example pins remain unchanged.

Calls ran October 2, 2026 on public SIE commit `23cc6df2eb6f86d23846459d7fcfa1dc4cec773c`. Model calls used remote GPU serving. This is a small functional illustration, not a new aggregate benchmark or measured hosted bill.

The visible hero uses IRS Publication 15-T (2026), issued December 3, 2025, page 20. The question is “What is the 2026 SEMIMONTHLY standard withholding for Married Filing Jointly with adjusted wages of $1,000?”. Both SIE page-image embeddings and OpenAI text-embedding-3-large over embedded PDF text rank page 20 first among eight candidate pages. It is a functional source match, not a rival failure. The page is the returned result; the displayed crop does not claim the embedding service extracted or calculated the withholding amount.

The page visibly says SEMIMONTHLY Payroll Period. For wages at least $995 and below $1,010, the married-filing-jointly standard-withholding cell is $0. The two crops are human-selected source regions from that same returned page, not model attribution. Each opens the complete page image. Page dimensions and crop coordinates are in the generated display JSON.

- [Pinned selected source input](https://huggingface.co/datasets/superlinked/sie-task-evidence/resolve/74f4e0eb2676a95b625783932796edac12040644/task-visual-examples-2026/irs-p15t-page-20.png), SHA-256 `273368aba69c5fc13f69ceb3de78876c6c97e2cf0069b6cd656695bd299a15a8`.
- `payroll-2026-p20.webp` SHA-256 `7dbb102faeb22f963d9f9de01e258f1b0388cd32c81d791c8c6f0f3acbd193dc`.
- Complete SIE rankings `visual-precise-search.json`, SHA-256 `631d3ec1172532e70a0379d479e2ad1cad2949b868e3f7e910b57ecd2f902f9e`; baseline rankings `openai-precise-search.json`, SHA-256 `4d6890fcdf7129073d1365ee464f37c19b0cc96f5277e7a2ce358dcb333fbf0b`.
- Model revision `bf790bd8780b098b86453444632a184bb770be1a`, compact 768-image-token cap.
- Generator: `apps/site/scripts/import-visual-document-search-illustration.mjs`, run against the downloaded packet and revision above.

All seven attempted questions and both methods' complete rankings are retained. The initial three generic questions omitted a wage amount and did not uniquely select an answer page. Four precise questions followed. Source inspection corrected the third initial expected page from 18 to 19; the original protocol remains unchanged. The first three precise questions favor the text baseline. The fourth was selected as an unambiguous current-edition page match readable on a phone, with both methods correct. `source-checks.json` records this selection and the manual source audit. These results do not replace or strengthen the separately pinned ViDoRe aggregate comparison. The illustration's baseline uses embedded text, while the study's baseline uses its supplied Markdown.

### Pain example

The pain section quotes seven words from
[docling-project/docling#4181](https://github.com/docling-project/docling/issues/4181), opened by bharm16 on
5 September 2026 against Docling 2.126.0, and links the complete report. It reports table text missing from both
Markdown and JSON exports, with only a log warning. The page does not claim that the native run tested that report's
document, or that page-image search fixes every parser failure.

### Run locally

The Run locally card links the [local setup recipe](foundations/LOCAL_SETUP.md). It pins the public server and SDK
source and the complete model snapshot the native run used.

## 3. Frozen studies and earlier page editions

Nothing in this part appears on the live page, and it supports no current claim. Its figures stay unchanged as the
record of an earlier edition of the page and of the runnable example. "That edition" below means that earlier edition.

### ViDoRe v3 study, 30 September 2026

An earlier edition of the page reported this study. It found that SIE's `tomoro-colqwen3-embed-4b`, served with its
`compact` profile, found the page that answers a question more often than OCR plus text embeddings, on every document
set tested, and ranked relevant pages higher than Voyage multimodal-3.5 and Cohere Embed v4 on nDCG@10. Against Cohere,
the right-page-first difference has a confidence interval that crosses zero, so no advantage is established on that
metric. Every figure in this section comes from one pre-registered run on 30 September 2026. The live page's quality
chart uses a different, later workload (part 2).

- **Recorded run:** the [`superlinked/sie-task-evidence`](https://huggingface.co/datasets/superlinked/sie-task-evidence)
  dataset at revision `45b85ef4efbbca1d2b4af7b8cbc8e83d5a354308`, folder `visual-document-search-vidore-v3/`. It holds the questions and their
  grades, every arm's top 100 pages for every question, and the recorded scores.
- **Runnable example:** [`examples/visual-document-search`](https://github.com/superlinked/sie/tree/6aa81068786a82c9a0d656f7a781e43a744400d6/examples/visual-document-search)
  in `superlinked/sie`. `score.py` recomputes every score below from the rankings and fails if one differs.
  `run.py` ranks the pages again on your own SIE server.

#### Benchmark

[ViDoRe v3](https://arxiv.org/abs/2601.08620) (CC BY 4.0) has human relevance grades (1 or 2) and a human evidence box
for every question. The run used six public datasets and only their English questions: 1,816 questions over 11,624
pages. Four datasets hold English documents: computer_science, finance_en, hr and pharmaceuticals. Two hold French
documents: energy and physics. On those two, every arm searched French pages with English questions. The study's
registration first called all six English-document datasets; its correction is published with the evidence
(`CORRECTION.md`), and the English-document results are below. Every page of a dataset is a candidate for each of its
questions. industrial was dropped for budget before any tested model saw it.

| Dataset | HuggingFace | Revision | Pages | English questions |
|---|---|---|---:|---:|
| computer_science | `vidore/vidore_v3_computer_science` | `d5cc75883d92e294f0c0fc2662551c9708a06ebc` | 1,360 | 215 |
| finance_en | `vidore/vidore_v3_finance_en` | `7f432c176d82e27546501ad8064a713ac3071809` | 2,942 | 309 |
| hr | `vidore/vidore_v3_hr` | `0cdf0979f2c5a0fd3e335e6373b9da48a9fe3bc3` | 1,110 | 318 |
| pharmaceuticals | `vidore/vidore_v3_pharmaceuticals` | `3abd4aa8a9445fb5538a78a19ba50bd57bd22b5c` | 2,313 | 364 |
| energy | `vidore/vidore_v3_energy` | `caec06d3c73434d635f710f93bcd898331c59f20` | 2,225 | 308 |
| physics | `vidore/vidore_v3_physics` | `a0de276f515acc044b72cae8de53a44bb5a8f1f5` | 1,674 | 302 |

Each page was rendered once to a JPEG with a 1,650-pixel long side (a US-letter page at 150 dpi), at quality 90. Every
image arm received those same bytes.

#### Arms

| Arm | Called as | Input |
|---|---|---|
| SIE tomoro-colqwen3-embed-4b:compact | open-source SIE server, `TomoroAI/tomoro-colqwen3-embed-4b:compact` (768 visual tokens per page), multivector, MaxSim | page image; question with `is_query` |
| SIE tomoro-colqwen3-embed-4b (default, 1,280 visual tokens) | the same model with its default profile | the same |
| Voyage multimodal-3.5 | `POST /v1/multimodalembeddings`, `input_type` document and query | page image; question |
| Cohere Embed v4 | `POST /v2/embed`, `embed-v4.0`, float, `search_document` and `search_query` | page image; question |
| Parse, then OpenAI text-embedding-3-large | `POST /v1/embeddings`, 3,072 dimensions, cosine | the page's ViDoRe markdown, truncated to 8,191 tokens; question |
| Parse, then BM25 | k1 1.2, b 0.75, lowercase alphanumeric terms, a fixed English stopword list | the page's ViDoRe markdown |

The two parse arms stand for the pipeline most teams run: parse each page to text, then search the text. ViDoRe's
markdown comes from a strong parser, so these arms are that pipeline at its best. The SIE arms ran on the open-source
SIE server on one H100, because SIE Cloud does not serve these models yet.

Gemini Embedding 2 was not run: no key was available. It lists $0.12 per 1,000 images real-time (a batch tier also exists), well below
every price in that edition, and its only published document-retrieval score is ViDoRe v2 (62.4 nDCG@5). That edition made no
claim about it, on price or on quality.

#### Results

Two metrics, each averaged per dataset and then over the six datasets. nDCG@10 uses the relevance grades as gains; it
is ViDoRe's own metric and the one the chart plots. "Right page first" means the first page returned has a positive
grade.

| Arm | nDCG@10 | Right page first |
|---|---|---|
| SIE tomoro-colqwen3-embed-4b:compact | 64.1 | 62.9% |
| SIE tomoro-colqwen3-embed-4b (default, 1,280 visual tokens) | 65.5 | 65.3% |
| Voyage multimodal-3.5 | 61.0 | 60.5% |
| Cohere Embed v4 | 61.1 | 60.8% |
| Parse, then OpenAI text-embedding-3-large | 56.0 | 54.4% |
| Parse, then BM25 | 42.0 | 41.3% |

Per dataset, nDCG@10 / right page first:

| Arm | computer_science | finance_en | hr | pharmaceuticals | energy | physics |
|---|---|---|---|---|---|---|
| SIE tomoro-colqwen3-embed-4b:compact | 77.1 / 81.4 | 65.7 / 64.4 | 62.2 / 61.3 | 67.8 / 70.1 | 63.1 / 57.8 | 48.7 / 42.4 |
| SIE tomoro-colqwen3-embed-4b (default, 1,280 visual tokens) | 76.4 / 80.5 | 69.2 / 70.9 | 63.5 / 63.8 | 68.5 / 69.2 | 65.3 / 60.4 | 49.9 / 47.0 |
| Voyage multimodal-3.5 | 73.9 / 80.0 | 61.3 / 63.4 | 57.7 / 56.3 | 64.5 / 62.4 | 62.2 / 57.5 | 46.4 / 43.7 |
| Cohere Embed v4 | 73.6 / 80.0 | 64.1 / 64.7 | 59.9 / 58.5 | 65.2 / 66.2 | 59.9 / 54.9 | 44.1 / 40.7 |
| Parse, then OpenAI text-embedding-3-large | 66.6 / 74.0 | 56.4 / 56.6 | 52.0 / 50.9 | 62.4 / 61.5 | 53.1 / 45.5 | 45.4 / 38.1 |
| Parse, then BM25 | 62.6 / 69.3 | 53.0 / 53.1 | 49.2 / 47.8 | 55.9 / 52.2 | 16.2 / 14.3 | 15.1 / 10.9 |

SIE `:compact` minus each arm. The intervals come from a dataset-stratified paired bootstrap: questions resampled within
each dataset, 10,000 draws, seed 20260930. The test on right page first is an exact McNemar.

| Arm | nDCG@10 | Right page first | Questions only SIE got right / only the arm did |
|---|---|---|---|
| SIE tomoro-colqwen3-embed-4b (default, 1,280 visual tokens) | -1.4 (-2.0 to -0.8) | -2.4 (-3.9 to -0.9) | 81 / 126, p = 0.0022 |
| Voyage multimodal-3.5 | +3.1 (+2.2 to +4.0) | +2.3 (+0.1 to +4.6) | 244 / 197, p = 0.028 |
| Cohere Embed v4 | +3.0 (+2.0 to +3.9) | +2.1 (-0.1 to +4.2) | 220 / 181, p = 0.058 |
| Parse, then OpenAI text-embedding-3-large | +8.1 (+6.9 to +9.3) | +8.5 (+6.0 to +11.0) | 344 / 189, p = 1.9e-11 |
| Parse, then BM25 | +22.1 (+20.7 to +23.4) | +21.6 (+19.1 to +24.1) | 522 / 124, p = 5.7e-59 |

##### English documents only

Over the four English-document datasets alone (1,206 questions), with the same bootstrap:

| Arm | nDCG@10 | Right page first | SIE `:compact` minus the arm, nDCG@10 |
|---|---|---|---|
| SIE tomoro-colqwen3-embed-4b:compact | 68.2 | 69.3% | |
| SIE tomoro-colqwen3-embed-4b (default, 1,280 visual tokens) | 69.4 | 71.1% | −1.2 (−1.9 to −0.6) |
| Voyage multimodal-3.5 | 64.3 | 65.5% | +3.8 (+2.7 to +5.0) |
| Cohere Embed v4 | 65.7 | 67.4% | +2.5 (+1.4 to +3.6) |
| Parse, then OpenAI text-embedding-3-large | 59.3 | 60.8% | +8.8 (+7.4 to +10.4) |
| Parse, then BM25 | 55.2 | 55.6% | +13.0 (+11.5 to +14.5) |

Every claim of that edition holds on these four as on all six.

BM25 was not on that edition's chart. It ran with an English stopword list, which fails on the French pages (16.2 and 15.1
on energy and physics), so its six-dataset score understates keyword search.

##### What that edition claimed, and why

- **More often than OCR plus text embeddings, on every document set.** The lower bounds are above zero on both metrics,
  and SIE leads in all six datasets on both.
- **More often than Voyage multimodal-3.5.** The pre-registered bar was a lower bound above zero on nDCG@10, and SIE
  leads in all six datasets on it. On right page first the macro lower bound is +0.1, and Voyage leads on physics
  (43.7% against 42.4%). That edition therefore stated this claim only on nDCG@10.
- **Cohere Embed v4:** SIE leads on nDCG@10 by 3.0 points (+2.0 to +3.9). On right page first the interval crosses
  zero, so that edition named Cohere only as a peer.
- **The default profile is 1.4 points better than `:compact`** on nDCG@10 (interval from -2.0 to -0.8), and 2.4
  points better on right page first. The pre-registration made `:compact` that edition's model because it matched the
  default on the pilot dataset. It clears every bar above, and it serves 2.2 times as many pages per second, which is
  what set that edition's price. The default profile is available too.
- **The test discriminates:** the best arm scores 65.5.

#### Prices

That edition priced one job: index a million consumed page images. Rival prices were read on 30 September 2026 and use their cheapest real-time tier, with no batch or asynchronous discount. SIE uses the selected commercial list price below. Rounding goes against SIE: its price is rounded up, and every rival's down. Query costs are excluded.

| Pipeline | Per 1,000 pages | A million pages | How |
|---|---:|---:|---|
| SIE tomoro-colqwen3-embed-4b:compact | $0.20 | $200 | Selected commercial list rate for consumed `:compact` images |
| Voyage multimodal-3.5 | $0.56 | $560 | $0.60 per billion pixels ([docs.voyageai.com/docs/pricing](https://docs.voyageai.com/docs/pricing)), at a 1,100-pixel long side (below) |
| Cohere Embed v4 | $1.14 | $1,140 | $0.47 per 1M image tokens (Azure's retail price list; cohere.com lists none); the run billed 28,386,276 image tokens for 11,632 pages, 2,440 a page |
| OCR + OpenAI embeddings | $1.57 | $1,570 | AWS Textract DetectDocumentText $1.50 per 1,000 pages for the first million ([aws.amazon.com/textract/pricing](https://aws.amazon.com/textract/pricing/); Azure Read and Google Document AI list the same), plus text-embedding-3-large at its $0.13 per 1M real-time rate on the recorded markdown, 607 tokens a page |
| OCR + keyword search | $1.50 | $1,500 | Textract alone |

The current indexing price requires `encode:TomoroAI/tomoro-colqwen3-embed-4b:compact:images:us`: exact USD `1 / 5000` per consumed image, $0.20 per 1,000 or $200 per million. It comes from the [commercial price list](/pricing/sie-cloud-prices.json), version `2026-10-10-production-bootstrap-v28`, effective `2026-10-10T18:00:00Z`, rate-book SHA256 `ab292600c88c4d95dfb67f7f568f9845eb738752208c9c4ff3bbd64ef3e9455d`. $0.20 is less than a seventh of the rounded-down $1.57 parse-then-embed price. A missing exact row fails the page-data build.

Text queries have a separate `encode:TomoroAI/tomoro-colqwen3-embed-4b:compact:input_tokens:us` rate: exact USD `1 / 500000` per actual processed input token, $2.00 per million. Both dimensions appear in the [model catalog](/cloud-models). Query image inputs are refused. The saved study does not record SIE's actual processed query-token usage, so it supports no SIE query total or cost per million questions. The compact visual-token ceiling and the caller's character count are not billed query-token usage.

These are selected list prices in a mixed book with `price_basis: selected-commercial-or-modelled-provider-cost-or-market-anchor` and `rate_grade_campaign: false`. They establish no measured serving cost or margin. The current list rate equals the study's original $0.20 per 1,000-image proposal, which was 2.2 times its recorded cost estimate; that estimate remains historical. At the 2026-09-30 study date the public main branch supported self-hosting the compact profile; hosted access and that proposal were unavailable. Importing a commercial price establishes no new hosted-availability observation.

##### Voyage at its cheapest resolution (computer_science)

Voyage bills a page by its pixels, capped at 2 million. The run scored three renders on computer_science:

| Long side | nDCG@10 |
|---|---:|
| 1,650 px | 73.9 |
| 1,100 px | 73.4 |
| 780 px | 72.6 |

1,100 pixels is the smallest render within a point of the 1,650-pixel score, so that edition priced Voyage there. Voyage's
quality everywhere else in that edition was its 1,650-pixel score, so Voyage gets its best quality at its lowest price.

##### Historical throughput behind the study proposal

The open-source SIE server on one H100 encoded 15.4 pages a second with `:compact` and 7.05 with the default profile.
Those figures are for 600 finance pages, msgpack responses, 8 and 16 requests in flight. An L4 encodes about 1.9 pages
a second with the default profile. Latency was not in that edition: at the 2026-09-30 study date SIE Cloud did not serve the model, so there was no hosted request to time.

#### That edition's examples

The historical study examples are questions where the OCR-plus-embeddings arm put a page with no relevance grade first, and SIE
put a grade-2 page first. The run has 231 such questions against that arm, and 100 the other way round (SIE wrong,
the arm right). Against Voyage the counts are 152 and 100, and against Cohere 136 and 95. The four historical illustrations were chosen by
eye among questions whose answer is in a table, a chart or a slide, from documents whose licence allows reuse. The
highlight on each returned page, and the note under it, say what that page shows instead of the answer. The rank is
where the arm put the page with the answer.

| Dataset and question id | Question | Where the text arm ranked the answer |
|---|---|---:|
| pharmaceuticals 44 | Identify the number of New Molecular Entities approved by CDER in fiscal year 2019 that received priority review. | 2 |
| finance_en 122 | How much did the amount of JP Morgan Chase's total asset management fee increase from 2023 to 2024? | 4 |
| hr 197 | What was Romania’s employment rate for people aged 15 to 24 in 2020? | 76 |
| hr 136 | Determine the year-on-year percentage growth in real wages recorded for the EU during the second quarter of 2024. | 2 |

Page renders. The first hash is of the JPEG the models received; the second is of the WebP the page shows.

| File | Dataset | Question | Corpus id | Document | Recorded JPEG SHA-256 | Displayed WebP SHA-256 |
|---|---|---:|---:|---|---|---|
| `fda-cder-2019-new-drugs-slide-18.webp` | pharmaceuticals | 44 | 1766 | CDER_New_Drugs_Program_2019_Update | `bc9359fa1887660aed14e3dffee908ba74036139d23fca33e639857370381ce2` | `2d0d97bad7dacbfd07873765ec15fa470dd357febf75f9108547df7626b5b46f` |
| `fda-cder-2019-new-drugs-slide-5.webp` | pharmaceuticals | 44 | 1784 | CDER_New_Drugs_Program_2019_Update | `2e0ed22534841da8d01eae7e27b42c98c6261e9aaa521e6b8dbd319a9dcd6e30` | `51fa5e104371e12fa15493ae091d5d761f11784c9af062294a5b6b443b7dc059` |
| `jpmorgan-2024-10k-p86.webp` | finance_en | 122 | 424 | jpmorgan_chase_2024 | `db9c1d10cb7c43609e2af580ae76e5d1a3f63bb21a243a1a296a99fbb43c64e9` | `c4372096117380b7be9eaa2c167bdfde84e38fca751109a4fe210a99208a5a40` |
| `eu-labour-market-2024-p16.webp` | hr | 197 | 751 | labour_market_and_wage_developments_in_europe-KEBN24001ENN | `05a8b17ca1f19ca430b1aae4b983dbc072f364b6d738692cc41634c9b4e20040` | `af53650d9ea599c9ce82116cca8e716af03f1bfd87b72073cc1a356c063ec118` |
| `eu-labour-market-2024-p52.webp` | hr | 136 | 791 | labour_market_and_wage_developments_in_europe-KEBN24001ENN | `e7b2d2771fa51984e2b257da86eb86737e35e88a6ffdc7caf67a361769082629` | `90d39aafd74f8bb2f40b8a47941f67b29294a126efd91020da44b3879bd521f2` |

Documents:

- `CDER_New_Drugs_Program_2019_Update`: https://www.fda.gov/about-fda/center-drug-evaluation-and-research-cder/meeting-presentations-drugs. Licence: Unless otherwise noted, FDA website text and graphics are public domain and may be republished without permission. Source credit is appreciated. [FDA website policies](https://www.fda.gov/about-fda/about-website/website-policies#linking).
- `jpmorgan_chase_2024`: https://www.sec.gov/Archives/edgar/data/0000019617/000001961725000270/0000019617-25-000270-index.html. Licence: Public Domain (Website section of https://www.sec.gov/about/privacy-information)
- `labour_market_and_wage_developments_in_europe-KEBN24001ENN`: https://op.europa.eu/fr/publication-detail/-/publication/057f23e9-bdc5-11ef-91ed-01aa75ed71a1/language-en. Licence: cc-by-4.0
- `joint_employment_report_2025-KE0125018ENN`: https://op.europa.eu/en/publication-detail/-/publication/b33bffec-e241-11ef-be2a-01aa75ed71a1/language-en. Licence: cc-by-4.0

#### Models not run

- tomoro-colqwen3-embed-8b scores higher on the public ViDoRe board. Its current remote code does not load under the
  server's pinned transformers 4.57, so it was dropped before the run.

### Earlier edition's closing diagram: index directly from page images

That edition's closing diagram added a different point from its quality bars: the application's indexing work. It renders a page, SIE encodes its image, and the application stores and scores page multivectors. The alternate route parses the page to Markdown before text embedding. Rendering, indexing and scoring remain the caller's work. No fabricated parser output or new parser failure is implied.

## Rights

Rights basis: the selected government table is public domain under 17 U.S.C. 105; [GovInfo policy](https://www.govinfo.gov/about/policies). The full IRS publication is excluded because other pages contain third-party photographs.
