mynkchaudhry/Florence-2-FT-DocVQA
Primitive: /extract · Extract ·
Florence-2
This model card provides details about the Florence-2-FT-DocVQA model, which is fine-tuned for Document Visual Question Answering (VQA) tasks.
Overview
Hardware: — drives latency, throughput & cost
| Size | 271M params |
|---|---|
| Tasks | /extract |
| License | apache-2.0 |
| Languages | en |
| Latency | 1.6 s |
| Throughput | 510 tok/s |
| Cost | $0.436 /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Extraction
| Output kinds | text_regions |
|---|---|
| Inputs | text · image |
| Max sequence length | — |
Benchmarks
DocVQA
Visual question answering on document images