# Structured output page sources

## Recorded model output

- Model: `Qwen/Qwen3.8-27B-FP8`
- Hugging Face revision in the SIE model config: `017b9c7af6b5689d5dd426a76e0bc077eb5ca20a`
- Served revision header (`X-Sie-Model-Revision`): `8bd714204e67a1c6c81f84b0dc486b6a6e96e943c42ff488f6b3cbf936e07955`
- Model card: <https://huggingface.co/Qwen/Qwen3.8-27B-FP8>
- Endpoint: `POST https://api.superlinked.com/v1/chat/completions`, server `0.7.3`
- Page run: 2026-09-15, from `21:02:44Z` to `21:04:24Z`, one request per document
- Playground run: 2026-09-15 at `22:27:37Z`, one request for GSA lot 377058, sent after its checks were committed
- Schema amendment and re-run: 2026-09-24, from `04:16:15Z`, when the first
  request was sent, to `04:18:50Z`, when the last response landed, eight
  requests. Both ends are read off the recorded calls: the first is the earliest
  `run_at` in the manifest and the second is the `Date` header of the last
  response. An earlier draft of this line gave a window typed rather than
  derived, and it disagreed with the manifest and with this file's own later
  statement of the earliest call. Two fields were dropped from the schemas and the eight documents
  carrying them were re-recorded; the section below says which and why. The
  served model revision and server version are the same on both dates, so the
  amendment did not cross a checkpoint.
- Runner: `apps/site/tests/fixtures/reference/structured-output/run.py`
- Display evidence generator: `apps/site/tests/fixtures/reference/structured-output/build_display.py`,
  which reads the fixtures and writes the page's evidence JSON. The file it
  replaces was written by a script that was never committed, so the generated
  artifact could not be regenerated. This one was checked by regenerating the
  pre-amendment evidence from the pre-amendment fixtures and comparing it with
  the committed artifact before it was trusted to produce a new one.
- Runnable example: <https://github.com/superlinked/sie/tree/a3edf835906814f9f5235fd0f7a47aaacee42dc6/examples/structured-output>

Every case sends the request the page's playground snippet shows:

- a system message with the document-type instruction, a blank line, `Required JSON schema:` and the schema printed with two-space indentation;
- a user message with the verbatim source text;
- `max_completion_tokens: 512`;
- `response_format: {"type": "json_schema", "json_schema": {"name": "structured_output", "strict": true, "schema": <schema>}}`;
- no sampling fields, so the model profile defaults apply.

Results across the 10 documents the page counts:

- 10 of 10 responses finished normally, parsed as JSON and validated against their schema with `jsonschema`;
- 84 of 85 acceptance checks passed. Every check was committed to `checks.json` before the run that scores it, and the two the amendment removed were removed together with the schema properties they scored. That ordering rests on this repository's own commit history, which is a record we keep rather than one a reader can hold us to; nothing in the recorded calls timestamps a check;
- median latency 8,804.75 ms. The page states that figure in its quality panel and makes no speed claim anywhere.

### Reproducing these figures without an API key

The example linked above fetches the recorded calls from the public Hugging Face
dataset and re-derives every figure this page publishes, offline, with no API
key and no inference spend. It prints 21 of 21 yes-or-no fields, 5 of 5 picks
from a fixed list, 58 of 59 everything else, 84 of 85 overall, and 0 documents
saying true or false, and exits nonzero if any of them fails to reproduce.

The link pins commit `a3edf835` rather than `main`, so a later push to the
example cannot change what this page cites. It moved there on 2026-09-24 from
`26c81c24`, the squash-merge of superlinked/sie#352, when every task page
re-pinned to the commit whose README links its page's `SOURCES.md`. Only that
README differs between the two: `fetch.py` and `score.py` are the same blobs at
both commits (`cc7ab7f5` and `009783e3`).
That example's own `fetch.py` pins a dataset revision as well,
`d04453560acaa308032c98feb8fe170d0b3b400e` of
[`superlinked/sie-task-evidence`](https://huggingface.co/datasets/superlinked/sie-task-evidence),
given in full here so the recorded calls can be fetched without opening the
example first. The two moved in one commit: an example whose constants matched this page while
its data did not would fail for every reader who ran it.

Checked by content on 2026-09-24 rather than by status code, fetched with no
auth. The run below was performed at `26c81c24`, which the link pinned at the
time. It carries to `a3edf835` because the files it exercised are byte-identical
at both commits, as the blob hashes above show; the README, the one file that
does differ, is not read by `fetch.py` or `score.py`. Said plainly rather than
restated as if it were re-run: this table records a run against `26c81c24`.

| Fetched | Result |
| --- | --- |
| `fetch.py`, `score.py` and `README.md` at the pin | 200 |
| a file that does not exist at that commit | 404 |
| the same SHA with one character changed | 404 |
| `fetch.py` at the pin, read | carries `REVISION = "d0445356…"` and its `MANIFEST_SHA256` together |
| `score.py` at the pin, read | `EXPECTED` opens `cases: 10` and `checks_passed: 84` |
| **both, downloaded into an empty directory and run** | `fetch.py` verified 5 files and 94 recorded calls, then `score.py` exited 0 on "84 of 85 checked fields right" |

The last row is the one that settles it. A status code proves nothing, because
an older commit of the same example still serves a complete, working demo of
whatever the page said then; and reading the constants proves only that the
strings look right. Running what the pin serves, in a directory holding nothing
else, is the check a reader performs.

Site CI asserts that the commit in `src/data/reference/task-examples.ts` and the
commit in this file are the same. It does not fetch the URL, so it cannot tell
you the example still reproduces; that check is the one above, run by hand when
the pin moves.

### How the 85 checks divide

The proof heading counts the fields a schema declared as a yes-or-no question,
meaning `"type": "boolean"` or `"type": ["boolean", "null"]`. There are 21 of
them across the 10 counted documents and all 21 are right. Sixteen are plain
booleans; five admit null, and two of those correctly came back null because the
GSA listing gives no answer. Not one of the ten source texts contains the word
true or false, which `structured-output.ts` and the page-data test both check
against the recorded texts rather than against this sentence.

All 85, by the kind of question the schema asked. The 21 above are the first
row; the other 64 are the two below it:

| What the schema asked for | Right |
| --- | --- |
| Yes-or-no fields (`boolean`, or `boolean` with null) | 21 of 21 |
| A pick from a fixed list (an enum, or an array of enum members) | 5 of 5 |
| Everything else: numbers, strings and dates | 58 of 59 |

The one miss over the counted set is `nhtsa-11443034.injured_people`, and no
surface of the page shows that document.

## The schema amendment, and what it cost

Two fields were removed from the schemas on 2026-09-24, and their acceptance
checks were removed with them. Neither was consistently grounded across the
document group whose schema asked it: one of the three GSA listings cannot
ground `condition`, and two of the five NHTSA narratives cannot reach their
coded `components`. A schema belongs to a genre rather than to a document, so
each field was dropped from its whole group rather than from the documents that
could not answer it. The tables below carry the per-document detail.

An earlier version of this sentence said both fields asked for a value the
document does not state, which is true of those three documents and false of
the other five.

**`condition`, on the three GSA lots.** The enum asks for a GSA condition grade,
and two of the three listings state theirs verbatim: lot 377054 says "This lot is
in scrap condition" and lot 377058 says "Usable/good condition." For those two
the field is perfectly answerable and the model answered it.

Lot 377053 states no grade at all. Its `agency_coded` is null, the registered
answer `unknown` is a human reading of that absence, and the listing does say
"Parts may be missing, and repairs may be required", so `repairable` is what its
words support. The enum has no member for "this listing does not say", because
`unknown` is being asked to mean both that and a stated unknown condition.

So the field was dropped from the GSA schemas **as a group**, not because no GSA
listing states a condition, but because one of the three cannot ground it and a
schema belongs to a document genre rather than to a document. Dropping it for
lot 377053 alone would have meant two GSA schemas differing by one field, and a
page that reported a figure over a set whose members were asked different
questions.

**`components`, on the five NHTSA complaints.** The enum is NHTSA's own component
taxonomy and the registered answers come from each complaint's
`agency_coded.components`, not from the narrative. How far the narrative reaches
the coded value varies by complaint, and the honest account is per complaint
rather than one sentence over five:

| Complaint | NHTSA coded | Reachable from the narrative? |
| --- | --- | --- |
| 11180068 | `STEERING` | Yes. "THE STEERING SEIZED" |
| 11532050 | `FIRERELATED` | Yes. The vehicle caught fire |
| 11443034 | `VEHICLE SPEED CONTROL`, `SERVICE BRAKES` | Partly. The narrative has the brake pedal and unintended acceleration |
| 11231274 | `AIR BAGS` | No. The narrative carries both air bags and a fire, and which one NHTSA codes is the convention |
| 11427180 | `ELECTRICAL SYSTEM` | No. The narrative never says electrical; it says the vehicle caught fire |

Two of the five cannot be reached from the words. The field survived four of
five partly because the pre-registered checks widened the target where the coded
value alone was unreachable: 11427180's check is `contains_any` over
`ELECTRICAL SYSTEM` and `FIRE RELATED`, and it passed on `FIRE RELATED`.

So this field was dropped from the NHTSA schemas as a group for the same reason
`condition` was dropped from the GSA ones: a schema belongs to a document genre,
some documents in the genre can ground the field and some cannot, and asking the
members of one counted set different questions is worse than not asking this
one.

The amendment was committed in its own commit before any call, and that commit's
author date, `2026-09-24T04:16:08Z`, is seven seconds before the earliest amended
call at `04:16:15Z`. Its committer date is later, because the branch was rebased
onto the re-pin afterwards and a rebase rewrites committer dates; the author date
is the one that survives, and it is the same basis the first run's note uses.

Seven seconds is thin, and an author date is a value this repository sets, so it
is a record we keep rather than a thing a reader can check against us. That is
the honest weight of the timing claim and it is stated rather than propped up.

An earlier version of this paragraph propped it up. It said
`structured-output-fixtures.test.ts` rebuilds every request body from
`cases.json` and recomputes every check result from `checks.json`, so a schema or
a check edited after its run fails those comparisons whatever the timestamps say.
That is wrong. Those are **consistency** checks: they catch an edit to one side
of a pair, and an edit to `checks.json` together with the recorded result it
scores passes both. They establish that the files agree, not when any of them
was written.

What the recording itself settles is narrower and does not depend on us. Each
reply carries exactly the properties of the narrowed schema and no others, which
is what `response_format: {"type": "json_schema", "strict": true}` produces, so
the schema that was sent is the narrowed one. `gsa-377053` came back without
`condition` and `nhtsa-11231274` without `components`. Forging that would mean
forging a server reply, alongside its `X-Sie-Request-Id`, its `Date` and the
served deployment revision the response carries. It says which schema was sent.
It says nothing about when the checks were written, and neither does anything
else in the evidence.

No surviving check was edited and no source text moved. The two SEC filings carry
neither field and were not re-run.

### It cost a field that had been right

The amendment was predicted to give 85 of 85. It gave 84 of 85.
`nhtsa-11443034.injured_people` returned `1` on the first run and `0` on the
re-run. The narrative says the contact "had chest and neck pains due to hitting
the steering wheel and did go to urgent care but did not seek medical
attention", which contradicts itself in one sentence; the registered `1` comes
from NHTSA's `agency_coded` and the pains are stated, so `1` is the better
reading and `0` is a defensible reading of the second clause.

That field is answerable from this document, and from every other complaint that
carries it, which is what separates it from the two that were dropped: those
were answerable from some documents in their group and not from others. So this
is a model error, not a case for narrowing further, and it is published rather
than recovered. The run was not repeated to get a better sample.

### What the page displays

The page renders 5 of those 10 documents: the hero complaint, three proof cards
and the playground. The other five are recorded, counted in the totals above and
held in the page data, and no card draws them. Selecting five is a display
decision and removed no document, response or check.

The three cards are one per document genre and one per schema, chosen to carry
different arguments:

| Card | Recorded id | Why it is displayed |
| --- | --- | --- |
| Vehicle complaint | `nhtsa-11231274` | Two yes-or-no fields and a count of injured people from a list of five injuries |
| Officer appointment filing | `sec-sprouts-ceo` | Dates the appointment from a fiscal-year phrase and reads a promotion as internal, 10 of 10 right |
| Surplus vehicle listing | `gsa-377053` | Sets `mileage_conflict` by comparing two figures in different sentences and leaves two fields null |

Each card draws the same panel as the hero. Every case the grid draws returned
10 fields until the amendment; the NHTSA complaint returns 9 now that
`components` is gone and the SEC filing still returns 10, so the cards are one
row apart and the CSS stretches them to the tallest. The build fails if two
displayed cases ever come back more than one row apart, which was a strict
equality before the amendment moved it by exactly one.

That is a statement about the three cards and not about the counted set, which
is what the sentence it replaces claimed. The playground case never returned 10:
GSA lot 377058 was recorded against a six-field schema and returns five now.

Two of the three documents fit whole; the SEC filing is clamped with an
ellipsis, and each card states how many words the call was given.

No surface of the page renders a field the run got wrong, and
`src/data/reference/tasks/structured-output.ts` fails the build if the hero, a
card or the playground ever does. The rule used to be the opposite: the grid was
required to carry every recorded miss, so that a grid quietly becoming
all-passes could not read as the whole picture. What replaced it does that job
in one line under the grid, which states how many fields the run got wrong and
that the document carrying them is not on the page, the way `/detect` states the
two boxes that landed on the wrong object.

The page carries its figures beside the proof heading rather than under the
grid. Five sit there: 21 of 21 yes-or-no fields, 5 of 5 picks from a fixed list,
58 of 59 everything else, 8.8 s median per call, and 3 of 10 documents shown.
The three groups add back up on the page, 21 + 5 + 58 against the 84 and
21 + 5 + 59 against the 85, so a reader can reconstruct the decomposition rather
than take it on trust. Site CI compares all five against the built page and
asserts the arithmetic.

One line does follow the last card, and only one: the totals line stating the
single wrong field and that no surface shows the document carrying it. This
paragraph used to open by saying nothing followed the last card and close by
saying that line does, because the correction was added to the end instead of
to the sentence it falsified.

The playground uses a shorter case so its code and output read side by side:
the Condition & Markings section of GSA lot 377058, recorded against a six-field
schema and now sent a five-field one, because `condition` was one of the two the
amendment dropped. It was added after the page run. Its four remaining checks
were committed before the run that scores them, which passed all four.
`cosmetic_issues` has no check because the note gives no single right wording, so
it is schema-validated only. The playground
prints the schema one property per line; the snippet parses it and sends the
same schema object and system text as the recorded request, and site CI
compares the snippet's request body with `requests/gsa-377058.json`. GSA lot
377054 is counted in the totals and drawn nowhere.

The hero complaint (NHTSA ODI 11180068) is a single recorded sample, and the
page makes no stability claim.

The response supplies the JSON text, `finish_reason` and token usage. The API
returns no per-field confidence or source offsets, so the page draws no model
attribution. Numbered marks on the hero complaint are human-placed pointers to
the words that settle each field, and their chips are deliberately neutral
rather than accent-coloured. The accent used to mean one thing on this page, a
returned value disagreeing with its recorded expectation, and no surface renders
one now, so the marks stay neutral to keep the accent free rather than to avoid
colliding with a treatment that still appears. The page shows the returned JSON
as parsed.

A field key is a name, so it wraps only at its own boundaries. The record column
splits each key after an underscore or a dot and puts a `<wbr>` at each join,
the same rule `/ocr` and `/doc-field-extraction` use, and the CSS holds
`overflow-wrap: normal` on the key so an inherited `anywhere` cannot break a
name mid-identifier. `structured-output-page.test.ts` checks every recorded key
on every surface, not only the displayed ones.

The website does not serve the raw responses. Site CI compares the page data
with the non-served fixtures in `apps/site/tests/fixtures/reference/structured-output/`.

## Recorded but not shown

Three CPSC recall notices were collected and recorded with the same call:
recalls 25353, 25342 and 25343. They are not shown and are not counted in the
page totals because /sparse-embeddings already uses CPSC recall notices. All
three passed every acceptance check (30 of 30). Their requests, responses and
check results stay in the fixtures.

The fixtures also hold `diagnostics/` and `verify/`, which are recorded calls
under different requests. No figure on the page rests on either, and neither is
displayed.

## Primary sources

NHTSA and GSA records are works of the U.S. government and are in the public
domain under 17 U.S.C. 105. Do not imply agency endorsement. SEC filings are
public company documents; each company retains copyright, and the page quotes
short attributed excerpts.

### NHTSA vehicle complaints

- Publisher: National Highway Traffic Safety Administration, complaints API
- Downloaded: 2026-09-15
- Derivation: the complaint `summary` field, verbatim
- ODI 11180068, 2016 Kia Sportage:
  <https://api.nhtsa.gov/complaints/complaintsByVehicle?make=kia&model=sportage&modelYear=2016>
- ODI 11231274, 2019 Honda CR-V:
  <https://api.nhtsa.gov/complaints/complaintsByVehicle?make=honda&model=cr-v&modelYear=2019>
- ODI 11427180, 2019 Chevrolet Bolt EV:
  <https://api.nhtsa.gov/complaints/complaintsByVehicle?make=chevrolet&model=bolt%20ev&modelYear=2019>
- ODI 11532050, 2020 Subaru Outback:
  <https://api.nhtsa.gov/complaints/complaintsByVehicle?make=subaru&model=outback&modelYear=2020>
- ODI 11443034, 2019 Toyota RAV4:
  <https://api.nhtsa.gov/complaints/complaintsByVehicle?make=toyota&model=rav4&modelYear=2019>
- Downloaded file SHA-256 digests: listed per case in `cases.json`

### SEC Form 8-K officer appointments

- Publisher: SEC EDGAR; documents filed by the named companies
- Downloaded: 2026-09-15
- Derivation: primary document converted to text with whitespace normalized; the excerpt starts at the Item 5.02 heading and ends before the named marker sentence
- Sprouts Farmers Market, Inc., accession `0001575515-26-000041`, filed 2026-09-01; excerpt ends before "There are no family relationships":
  <https://www.sec.gov/Archives/edgar/data/1575515/000157551526000041/sfm-20260831.htm>
- Downloaded file SHA-256: `24903cedaf931c22c8e3bc302834e626cb08e2692b32b710b8426071e2b1ba4b`
- Magnachip Semiconductor Corporation, accession `0001193125-26-289285`, filed 2026-06-30; excerpt ends before "Additionally, the Company expects":
  <https://www.sec.gov/Archives/edgar/data/1325702/000119312526289285/mx-20260630.htm>
- Downloaded file SHA-256: `785ebbf328946b622542040daa973b656a49d0fb205e9ee542ea5ac87400d83a`

### GSA Auctions surplus vehicle listings

- Publisher: U.S. General Services Administration, GSA Auctions API, for the National Park Service (lots 377054 and 377053) and the Fish and Wildlife Service (lot 377058)
- API: <https://api.gsa.gov/assets/gsaauctions/v2/auctions?format=JSON>
- Snapshot downloaded: 2026-09-15, SHA-256 `e017efe26b638aef059da4a511182b38291a4c34ce65224c7d67fd82ded3358c`
- Derivation: the lot's `lotInfo` HTML converted to plain text, one line per heading or list item, wording verbatim
- Lot 377054, 2012 Dodge RAM, sale `3-1-QSC-I-26-583` lot 002, auction 2026-09-09 to 2026-09-16:
  <https://www.gsaauctions.gov/auctions/preview/377054>
- Lot 377053, 1984 IHC 1754, sale `3-1-QSC-I-26-583` lot 001, auction 2026-09-09 to 2026-09-16:
  <https://www.gsaauctions.gov/auctions/preview/377053>
- Lot 377058, 2017 Dodge Ram, sale `3-1-QSC-I-26-583` lot 006, auction 2026-09-09 to 2026-09-16; the case uses the contiguous Condition & Markings section, from its heading through "Additional deficiencies/information unknown.":
  <https://www.gsaauctions.gov/auctions/preview/377058>
- Active listings close when their auction ends; the snapshot file and its digest are the durable record.

### CPSC recall notices (recorded, not shown)

- Publisher: U.S. Consumer Product Safety Commission, Recalls REST API; U.S. government work, public domain
- API query: <https://www.saferproducts.gov/RestWebServices/Recall?format=json&RecallDateStart=2025-01-01&RecallDateEnd=2025-06-30>
- Recalls 25353 (Coleman Converta camping cots), 25342 (AstroAI minifridges) and 25343 (Polaris Ranger XP Kinetic ROVs); file SHA-256 listed in `cases.json`
