# Typed decisions: several scored questions about one record

This page demonstrates a named-question interface: one record is sent with
several choice groups, and one SIE request returns a distribution for each.
Caller code maps the two-option check to a boolean. It does not claim that
all answers are reliable enough for unattended vulnerability triage.

## Public evidence and runnable example

The frozen evidence is anonymously readable at
[superlinked/sie-task-evidence, revision bcf30acf1f1b18ae550715c42eefcb0c50783578](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/bcf30acf1f1b18ae550715c42eefcb0c50783578/typed-decisions).
Its manifest SHA-256 is
`b06b4b65afaf6dc281ffee18e147db75f36775cf66a136d3b18b575ace51cde5`.
The manifest checks all inputs, 22,384 recorded calls, summary and page surfaces.
The public scorer reproduces the published surfaces exactly.

The page pins the merged public example at
[8b4dc459dea70ca9bf9ff2320a459663b88ba2a3](https://github.com/superlinked/sie/tree/8b4dc459dea70ca9bf9ff2320a459663b88ba2a3/examples/typed-decisions),
the squash commit of [example PR #376](https://github.com/superlinked/sie/pull/376).
README, fetcher, scorer and runner were checked anonymously; a nonexistent
file at the same commit returned 404. Its `PREREGISTRATION.md` specifies the
questions, date splits, model settings, selection and withdrawal rules before
the held-out calls. No new inference was run for this page reconciliation.

## What was measured

The bespoke vulnerability-triage test has **160 test records**, 20 for each of
eight weakness classes, published by NVD in December 2023. Settings were fixed
on 224 earlier October/November records. Gold answers use the NVD analyst's CWE
and CVSS v3.1 assessment. The source record may contain information beyond its
description, so a correct answer is an inference judged against that assessment,
not an extracted quotation.

The displayed model is `knowledgator/gliclass-large-v1.0`, backend
`gliclass-large-v1-one-call`. Every question is a separate label group inside
one API request. The server encodes each group separately; one request does
not mean one encoder pass. Each record ran once, with no repeatability claim.

| Question | Exact result | Displayed accuracy |
|---|---|---|
| Weakness, one of eight classes | 142 of 160 | 88.8% |
| Attack vector, one of four classes | 147 of 160 | 91.9% |
| Remote exploitation without login, yes/no | 115 of 160 | 71.9% |
| Every question correct on the same record | 96 of 160 | 60.0% |

The last row is recomputed from the intersection of all three correct-answer
sets after the public scorer's frozen per-question rules. It is not a mean of
the rows above. The two choice questions together produced 289 correct answers
out of 320; that is an answer-level metric, not whole-record reliability.

The displayed scores are the recorded, normalized probabilities of the
model's own top options, as shown by the public example. Aggregate accuracy
uses the frozen per-question decision rules, which may adjust distributions. They are not a guarantee of
correctness. The scorer reports Brier scores and calibration error; a workflow
must validate its own decision and review thresholds. No statement that "90%
sure means 90% correct" carries over from the old classification study.

The model passed the pre-registered page rule for all three questions on dev
and retained them on test. This does not establish parity with an LLM. The
recorded `Qwen/Qwen3-4B-Instruct-2507` was stronger on attack vector (150 of 160)
and the yes/no check (129 of 160); its weakness answer failed the dev page rule.
The three GLiNER2.5-Decide variants did not pass the yes/no rule. The complete
model catalog, failed questions and withdrawn results remain in the public
example and evidence; no clean illustrated record overrides these aggregates.

## Recorded timing, scope and access

On the September 25 recording, the displayed model's median round trip was
27.795 ms, rounded once to **28 ms** in the generated page data. The server ran public commit
`d41ba7fb828cedb6a08f4c4028c6ce621385168f`, on one NVIDIA L4 with 10 CPU cores
and 40 GiB memory reserved. The client ran on the same machine, serially, one
record at a time, at `http://localhost:8080`. The run measures that self-hosted
setup. It is not current Cloud latency, throughput, cold-start time or an SLA.
The public code also supports a hosted base URL such as
`https://api.superlinked.com`; no hosted measurement is claimed here.

The recorded small Qwen LLM took 2,676.27 ms median for its call. The public
example's registered speed comparison passes, but this page does not pitch
parity with that model or a speed advantage over a current flagship. No price
comparison is published. The deployment cards offer the exact recorded raw-Large
model through self-hosting and a public released NVIDIA image. Public main supports model serving and label groups;
self-hosting cost depends on hardware and utilization. The trial buttons point
to the immutable example and public-main setup, with no hosted access promise.

## Illustrated records

The public `page.py` selection rule chose one record for each distinction:

- `CVE-2023-6826`: network reachable, but exploiting the upload flaw
  needs an account. The named answers are unrestricted file upload, network,
  and false for remote exploitation without login.
- `CVE-2023-49378`: cross-site request forgery, network, and true for the check.
- `CVE-2023-44278`: path traversal, local access, and false for the check.

Every displayed answer is correct against its record's gold and is the model's
own top option. The complete source descriptions are retained. No probabilities
or model output are authored. The primary workflow illustration is the recorded
2026 invoice described below. The archived `CVE-2023-6826` description remains visible beside its two named
answers, with the account-access condition highlighted. The other selected
records remain in the immutable packet. Their original dates, raw outputs and
scores are unchanged.

## Regeneration and checks

Fetch the pinned public example's seven evidence files, then run its standard
library scorer. `page.py` checks every figure it publishes against its frozen
assertions and writes the independently scored `page.json` fixture.

From a clean website checkout, regenerate the page with:

```sh
python apps/site/scripts/import-typed-decisions.py \
  /path/to/downloaded-public-evidence \
  /path/to/public-sie/examples/typed-decisions
```

The importer verifies the anchored manifest and all file hashes, pins the
public scorer/source-file digests, regenerates every public page surface and
checks equality with the published `page.json`. It also checks the complete
160-record one-request cohort, recomputes the all-questions intersection and
reconstructs every illustrated request against its recorded request digest.
The website tests compare generated figures with the independent public
scorer fixture and check illustrated text, answers, scores and example pins.

## Previous page and distinct jobs

The previous page presented a DeBERTa classifier followed by GPT-6 Sol over
12 single-label datasets. That is a classification/cascade study, not evidence
for multi-question record decisions. Its complete public recordings remain at
[the original common12 revision](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/adf623a3a6a8051e8cd8c96c8cf916d80dd6213f/typed-decisions/common12).
No figures from it are relabelled as this CVE study.

Classification returns label scores for a text. Typed decisions collects
several named answers about one record. Request routing selects the next
handler. This page does not change the separate classification or routing
models, measurements, routes or prices.

## Source and licenses

The public example records its NVD source URLs, retrieved-field hashes and
input construction. This product uses data from the NVD API but is not endorsed
or certified by the NVD. CVE descriptions use the
[CVE Program Terms of Use](https://www.cve.org/Legal/TermsOfUse), whose copyright
and license are retained in the frozen manifest, generated evidence and page.

Copyright (c) 1999-2026, The MITRE Corporation. CVE is a trademark and the CVE
logo is a registered trademark of The MITRE Corporation.

MITRE hereby grants you a perpetual, worldwide, non-exclusive, no-charge,
royalty-free, irrevocable copyright license to reproduce, prepare derivative
works of, publicly display, publicly perform, sublicense, and distribute Common
Vulnerabilities and Exposures (CVE). Any copy you make for such purposes is
authorized provided that you reproduce MITRE's copyright designation and this
license in any such copy. (CVE Program Terms of Use,
https://www.cve.org/Legal/TermsOfUse)

The tested open model is Apache-2.0. The workflow benchmark's separate 400-case
set, training-split fine-tuning caveat and third-party terms remain documented
in the example; that set supplies no page quality claim.

## Recorded invoice workflow illustration

The invoice illustration uses `invoice_processing_000038` from the separately recorded 400-case workflows set at the same immutable dataset revision `bcf30acf1f1b18ae550715c42eefcb0c50783578`. The source is LocalLLaMA/typed-decisions, revision `c76749ec58bd8c3d2ea706b31c333a9059c38f90`, Apache 2.0; its labels are teacher-model estimates, not human invoice decisions. This is an illustration, with no invoice accuracy, financial-safety or latency claim. The 160-CVE results remain a separate cohort.

`import-typed-decisions.py` verifies the existing manifest and source hashes, rebuilds the recorded one-call GLiClass large v1.0 request and checks its request digest. It exports two of that call's five questions: `matches_order=true` (99.8% model score) and `duplicate=true` (94.6%). Both match the recorded source labels. The full call also asked disposition, discrepancy severity and urgency; those fields are not displayed and do not authorize payment. Their original answers remain in the public recording. The full request is stored in the generated page evidence.

The input invoice `INV-2026-6576` is in its own prior-invoice list, and its $11,087 total matches the purchase order. These visible facts are asserted by the generator. The image-free input panel is a projection of the recorded state, not an uploaded invoice scan. The review queue is an authored example of a caller policy (`duplicate=yes` sends the case to review), not a recorded action or an additional model output. Model scores remain uncalibrated.
