# Knowledge graph page sources

/knowledge-graph shows what graph extraction returns: typed entities with their
character offsets, and directed relations between them, for entity and relation
types named in each request. It makes no comparison with other models or
services. The studies behind that decision are listed at the end of this file
with their published results.

## The model and the recorded calls

- Model: `fastino/gliner2-large-v1`, Hugging Face revision
  `b122b11eeaee4dabd32bed80412f3234c0d0e943`
- Served deployment revision, from the `X-SIE-Model-Revision` header of every
  recorded call: `10333b84de80b402376b626eb25366fb081d3faeb893eb4b01cf32e8c27e4aff`
- Endpoint `https://api.superlinked.com`, server version 0.7.3, recorded on
  15 September 2026
- Model card: <https://huggingface.co/fastino/gliner2-large-v1>

Each paragraph goes through the two `/v1/extract` calls the page's snippet
shows. The first sends the paragraph and its entity labels. The second sends
the same paragraph, the first call's entities as item metadata, and the
relation labels. No threshold is sent, so the server default of 0.5 applies.
SIE returns a relation only when both of its endpoints are among the entities
the second call received.

The full recording is pinned at
<https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/492389541d278e8b75cc9407e5170566e4523fa5/knowledge-graph>.
`canonical-calls.json` (SHA-256
`8426de61d623916217209ab3191fcd1a3fd63394064bd0574c87d30f6704715d`) holds forty
calls: ten 2026 paragraphs, GLiNER2 Base and Large, entities then relations,
with requests, responses, headers and per-entry digests. The runnable example
<https://github.com/superlinked/sie/tree/ef90dd602978cd7b02fe8820473e83e0f0f88191/examples/knowledge-graph>
verifies its Large-model calls against that recording. Site CI checks every
displayed entity, offset and relation against the recorded responses.

## The hero: Veracyte Form 8-K, 10 September 2026, Item 1.01

- <https://www.sec.gov/Archives/edgar/data/0001384101/000138410126000049/vcyt-20260910.htm>,
  a public SEC EDGAR filing quoted as a short attributed excerpt
- Paragraph SHA-256 `77fa25cb079f403f8b9a026e0f157eb23a093434bddccad8dcc0b4694aac59d4`
- Entity labels `company`, `subsidiary`, `product`, `disease`, `date`;
  relation labels `acquired`, `subsidiary of`, `develops`, `focused on`
- Recorded calls `veracyte-convergent__gliner2-large-v1__entities`
  (entry `3be98522e05acbf74436519166f5f4c686be9df7cbdaf18be7546e1c6c3edbfd`) and
  `__relations` (entry `59de945b3f1288cb356afd56badba8a6cad5f5b58bc32395ad0cb06fe6c6ac21`),
  93 input tokens each

The two sentences were read against the filing on 8 October 2026 and match it
verbatim. Each of the five returned edges is stated there: Convergent continued
as "a wholly owned subsidiary of Veracyte" (subsidiary of); the merger is "the
acquisition" of Convergent (Veracyte acquired Convergent); Convergent is "a
genomic diagnostics company focused on bladder cancer" (focused on); and the
acquisition adds "Convergent's UroAmp and proprietary urine tumor DNA
technology" (Convergent develops each). The page underlines every returned
entity where the text says it and places each edge under the sentence that
holds both of its endpoints.

## The API example: Flex Ltd. Form 8-K, 30 April 2026, Item 1.01

- <https://www.sec.gov/Archives/edgar/data/0000866374/000110465926054529/tm2612613d1_8k.htm>,
  filed 4 May 2026, a public SEC EDGAR filing quoted as a short attributed excerpt
- Paragraph SHA-256 `c640a11b9f98bd9a19f195cfdab16263e65657cae190b80e21e757ac2a5d94c3`
- Entity labels `company`, `bank`, `agreement`, `amount`, `date`; relation
  labels `borrower under`, `commitment amount`
- Recorded calls `flex-credit-facility__gliner2-large-v1__entities`
  (entry `d38da06cea30aaa9fa88c4873d3788d611d81918057b4d80ea1ce20b1918d57b`) and
  `__relations` (entry `70d3b59fb2c955c79007a24fa6254ff439c3fe14bf673b4b53a4ba9493d39e77`),
  92 input tokens each

The paragraph was read against the filing on 8 October 2026 and matches it
verbatim. Both requested relation types came back and both are stated: Flex
entered the Credit Agreement "as borrower" (borrower under), and the credit
facility has "an aggregate commitment amount of $1.45 billion" (commitment
amount). The API output lists those two relations with their native scores, the
call-one label and offsets of each endpoint, and the two entities no relation
used. Joining a relation endpoint to its entity span by text is the caller's
step; the page marks it as such.

Copy and Run use the registered run preset
`knowledge-graph-flex-credit-gliner2-large-2026-09-15`
(`packages/tasks/src/recorded-inputs/knowledge-graph-flex-credit.json`): the
same paragraph and labels on GLiNER2 Large, default profile. The hero's
sign-up opens the same preset.

## Deployment figures

The deployment card prices a paragraph graph from the two recorded Flex calls at
the published US rate for GLiNER2 Large, $0.18565866666 per million input tokens
(rate book `2026-10-02-production-bootstrap-v19`). Each 92-token call rounds up
to two credits, so one graph takes four credits; the $50 credit pack (5,000,000
credits) covers 1.2 million paragraphs of that length. Longer paragraphs cost
more.

## The studies, and why the page makes no comparison

Before choosing this form, we measured graph extraction against hosted language
models on Re-DocRED (MIT annotations, English Wikipedia text under CC BY-SA),
scoring typed, directed edges against human gold.

- Frozen inputs, gold and scorer:
  <https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/2cf6442b94ad12e72d59c3456337abd2a1943d48/knowledge-graph-pilot/2026-10-06>
- 100-article pilot results:
  <https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/607f2889b888b3d7deed5bd27ff3c691fc18888d/evaluations/2026-10-08/knowledge-graph-original100-results-v1>.
  Mean article F1: SIE Qwen3.8 27B 0.197, GPT-6 Luna 0.188, Claude Sonnet 5.5
  0.310. The SIE minus Luna difference, +0.009 (95% interval -0.025 to +0.040),
  did not meet the pre-declared gate. The GLiNER2 arm could not be served for
  that run and was recorded as unavailable, not scored.
- 40-article development screen, inputs and runner:
  <https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/0abe19c4683a7cb85b0b3a4b1951ded0f8a9137b/knowledge-graph-r1/2026-10-08>;
  results:
  <https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/55121207e38d899112e096094f03b0290a07984b/knowledge-graph-r1-results/2026-10-08>.
  Mean article F1: Qwen3.8 27B with the baseline prompt 0.251, with relation
  definitions 0.271, relation-first 0.178; Claude Sonnet 5.5 0.352. Every SIE
  recipe cost at most half of Sonnet per article and none came within the
  frozen 0.03 quality margin. Qwen3.5 122B did not finish loading within the
  bounded runs and has no score.
- Blinded review of saved false positives from the pilot:
  <https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/aafc7f9deec7ccd5ef2536a115e205d32c0c0be8/knowledge-graph-r0-blind-diagnostic/2026-10-08>.
  Of 50 sampled false positives per model, the article supported 18 for Qwen3.8,
  17 for Luna and 36 for Sonnet. Re-DocRED gold misses many relations the
  articles state, so these F1 scores understate every model; the review is
  qualitative and does not change any score.

The page therefore describes the native capability and its exact API, and
claims nothing about quality or price against other models.

## Source text

Each displayed paragraph is a contiguous span of its filing, with HTML tags
removed, entities unescaped and whitespace collapsed. Provenance, download
digests and the public-use basis of all ten recorded paragraphs are in
`inputs/candidates.json` of the pinned recording. This page claims no licence
over the filings themselves.
