---
title: "SIE Document Processing: Results from OCR and Field Extraction Tests"
description: We tested document-to-Markdown, OCR, and typed field extraction across reports, forms, invoices, labels, and technical documents. These are the results, including the failures.
canonical_url: https://superlinked.com/blog/sie-document-processing-ocr-field-extraction-tests
last_updated: 2026-10-05
---

A clean document-processing demo ends when text appears on screen. Production keeps going.

The table needs its numbers in the right cells. A two-column paper needs the right reading order. Dates printed in regional formats need a defensible interpretation. And if an agent extracts a field, someone must be able to trace it back to the page.

We recorded three document workflows in SIE Cloud:

- Converting complete documents into Markdown
- Reading page images once, then filling typed schemas from the recorded text
- Extracting fields from varied forms, with a second pass to correct the first result

These tests measure different things. The olmOCR score measures page reconstruction. Field accuracy measures extraction. The recorded checks look for exact headings, figures, order, tables, or schema values.

The numbers should stay separate.

<BlogSieCta />

## First job: preserve enough of the page

A document parser feeds everything downstream. If it scrambles a two-column paper, drops a heading, or moves a number into the wrong table cell, search and generation inherit the error.

We tested [SIE's document-to-Markdown task](/doc-to-markdown) on four document types:

- NVIDIA CFO commentary containing financial tables
- The two-column Docling technical report
- A SiriusPoint investor presentation
- A FEMA proof-of-loss form

The run passed 25 of 25 defined checks. Eighteen checks looked for a named heading or figure, three checked section order, and four counted tables.

The two-column paper came back in reading order. The NVIDIA outlook section returned all nine checked lines. Every figure checked in the SiriusPoint highlights table appeared in its own cell.

There were failures.

The SiriusPoint table repeated one header and returned two columns with the same Q1 '24 label. Farther down, those columns contained different figures. A reader could see the numbers, but the Markdown did not preserve enough header structure to tell which period each number belonged to.

The FEMA form returned all 12 amount labels and all six claim-type choices. It also produced 26 bare dollar signs without their field labels. The checkboxes landed 32 lines below the question they answered.

That distinction matters. Passing content checks does not mean every relationship on the page survived.

The [document-to-Markdown](/doc-to-markdown) page reports SIE LightOnOCR-2-1B at 0.752707 Overall on Ai2's olmOCR-Bench, equations included, and $2 per 1,000 pages. Those figures describe the published benchmark panel. The four-document run provides the more useful detail: which parts survived and where the structure weakened.

## Second job: read the image once

Many document systems reopen the original image every time the application needs a different set of fields. That repeats the expensive visual step and makes each result harder to compare.

The [OCR workflow](/ocr) splits the work.

LightOnOCR 2 1B reads a page into Markdown. A generation model then reads that saved text and fills a strict JSON schema. Add another schema and the system reuses the Markdown; it does not send the image through OCR again.

The recorded run, dated September 16, 2026, covered six documents. Three are shown on the task page.

One was a photographed product label. Its printed date, `12/02/2024`, could mean February 12 or December 2. The OCR output also retained the GS1 string under the barcode. Application identifier 15 contained `240212`, so the typed result returned `2024-02-12`.

All seven registered fields matched the label.

Another page contained financial values printed in parentheses. The typed schema required integers, so `(259)` needed to become `-259`. All 25 checked values on that report page were returned with the expected signs, and the reconciliation total remained positive.

A rendered invoice added more noise. Its title, software licence, and repository link appeared around the invoice itself. OCR read the whole page. The schema selected the invoice fields and left the surrounding page furniture out of the record. All 13 checked fields matched.

Across the three schemas shown, 45 of 45 fields matched the printed pages. Across the full six-document run, seven of seven replies parsed and validated against their schemas.

That LightOnOCR page-reconstruction score lives on the document-to-Markdown task page. The [OCR](/ocr) page measures a different job: turning receipt and handwriting photos into text with Z.ai GLM-OCR. Do not treat those panels as one result.

Pick the model according to the page and the consequence of an error.

## Third job: correct fields without breaking good ones

Field extraction gets harder when layouts change.

A fixed invoice template can rely on coordinates. A production queue might contain an aviation certificate, an electricity bill in Greek, a phone photo of a receipt, and a scanned workplace injury log before lunch.

The [document field-extraction test](/doc-field-extraction) used eight documents across seven layouts. The first pass returned 180 of 223 checked fields exactly. A second text-only pass raised that result to 204 of 223.

Zero fields that were correct after the first call became wrong after the second.

On a NIST certificate, the correction pass separated all 14 element names from their chemical symbols. On the Greek electricity bill, it corrected four current and previous meter readings that the first pass had reversed. On an OSHA injury log, it corrected two fields while keeping five cases as five separate typed objects.

The correction pass has a cost. It reads no image, but the task page says it costs about the same as the first call. The shortest document took 24 seconds over both calls for 11 fields. The longest took 135 seconds for 71.

Use it where the extra accuracy earns its place. A low-value intake form may not need two calls. A certificate, claim, or ledger entry might.

The field-extraction page reports SIE Qwen3.8 27B FP8 JSON accuracy from a 700-business-document study. That benchmark and the eight-document correction run answer different questions.

## What these document steps let an agent do

Document processing is rarely the final product. It prepares evidence for the next decision.

A [legal agent](/legal) can read a 42-page master agreement, combine it with a commercial schedule, client instructions, and firm precedent, then retrieve the clause that changes a liability cap. OCR and document-to-Markdown establish the searchable record. Retrieval and structured output turn that record into a cited finding.

In [finance](/finance), the problem is often version control. The example agent compares an original Form 10-Q with a later Form 10-K/A and returns the controlling figure from the amended filing. That workflow depends on tables surviving conversion and every number retaining its source.

An [insurance agent](/insurance) has to read the appeal, the policy language, estimates, and prior-loss records together. The published example checks which debris-removal costs fall inside a flood policy. The answer points back to the policy and appeal record.

A [healthcare documentation agent](/healthcare) compares submitted records with current CMS criteria. In the published L1851 example, the record says the face-to-face encounter happened seven months earlier. The rule allows six. The useful output is a cited mismatch for a human reviewer.

The same pattern appears in [manufacturing](/manufacturing). A reliability agent reads an NTSB report and joins it with sensor events and alert-routing rules. In the example, it reconstructs a sequence of 38°F, 103°F, and 253°F above ambient, along with who received each alert.

Different industries. Same dependency: the source page must survive before the agent can reason over it.

## What the tests support

The evidence supports a practical claim.

SIE can convert varied documents into Markdown, reuse OCR output across schemas, and return typed fields from layouts it was not configured for one by one. The recorded runs expose their inputs, checks, outputs, and failures.

The tests do not support perfect table reconstruction. They do not remove human review from high-consequence decisions. And they do not make every document equally easy.

That honesty makes the results useful. You can see where a small OCR model is enough, where a correction pass helps, and where damaged structure should stop the pipeline for review.

Run the [document-to-Markdown](/doc-to-markdown), [OCR](/ocr), or [field-extraction](/doc-field-extraction) examples in [SIE Cloud](/cloud) and test them on the documents your agent actually receives.

## Frequently asked questions: AI document processing

### What is AI document processing?

AI document processing turns PDFs, scans, forms, images, and office files into information software can use. A typical workflow reads the page, preserves its text and structure, then extracts selected values into typed fields or searchable Markdown.

### What is the difference between OCR and document-to-Markdown?

OCR reads text from an image or scanned page. Document-to-Markdown also tries to retain headings, tables, lists, and reading order. That added structure helps agents search, cite, and reason over the document.

### Can SIE extract structured data from invoices and forms?

Yes. SIE can read a document and return fields that match a supplied JSON schema. The recorded field-extraction test covered eight documents across seven layouts and returned 204 of 223 checked fields exactly after a correction pass.

### How accurate is SIE document processing?

Accuracy depends on the task and document. The document-to-Markdown run passed 25 of 25 defined checks across four document types. The OCR workflow matched 45 of 45 checked fields across the three schemas shown. A separate field-extraction run reached 204 of 223 exact fields after two calls.

### Can SIE preserve tables in PDFs?

SIE can preserve table rows, cells, and figures, although complex headers can still lose context. In one recorded investor-presentation test, every checked figure stayed in its cell, while duplicated column labels made two periods ambiguous. High-consequence tables should still go through validation.

### Can the same OCR result fill several schemas?

Yes. SIE can store the Markdown from one OCR pass and run further schemas against that text. The original image does not need to be processed again for every new field set.

### When should a document workflow use a correction pass?

Use a correction pass when the value of fixing extraction errors exceeds the cost and added latency. The recorded test improved from 180 to 204 exact fields out of 223, with zero previously correct fields becoming wrong.

### Does AI document processing remove the need for human review?

No. Human review remains necessary when an error could affect a payment, claim, contract, medical decision, or safety investigation. The model should return source-linked evidence so the reviewer can inspect the relevant page.

### Which industries can use document-processing agents?

Legal teams can review agreements and precedent. Finance teams can compare original and amended filings. Insurance teams can check claims against policy language. Healthcare teams can test submitted records against current criteria. Manufacturing teams can connect reports, manuals, and maintenance evidence.

### Can SIE process private documents?

Yes. Teams can use [SIE Cloud](/cloud) or run the same inference stack in their own environment. Self-hosting keeps document processing within infrastructure the team controls, including [offline deployments](/docs/deployment/offline).
