naver-clova-ix/donut-base-finetuned-cord-v2
Primitive: /extract · Extract ·
Encoder-Decoder
Donut model fine-tuned on CORD. It was introduced in the paper OCR-free Document Understanding Transformer by Geewok et al. and first released in this repository.
MultimodalText regions
Overview
Hardware: — drives latency, throughput & cost
| Size | 110M params |
|---|---|
| Tasks | /extract |
| License | mit |
| Latency | 8.4 s |
| Throughput | 757 tok/s |
| Cost | $0.294 /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Extraction
| Output kinds | text_regions |
|---|---|
| Inputs | image |
| Max sequence length | — |
Benchmarks
CORD v2
Key information extraction from receipt images
Corpus: 100 Queries: 100
Quality
f1 0.1123
Performance L4 b1 c16
Compare (0)Compare →