Why did we open-source our inference engine? Read the post

← Catalog

mynkchaudhry/Florence-2-FT-DocVQA

Open comparison →

Primitive: /extract · Extract · Florence-2

This model card provides details about the Florence-2-FT-DocVQA model, which is fine-tuned for Document Visual Question Answering (VQA) tasks.

MultimodalText regions

Overview

Hardware: — drives latency, throughput & cost

Size271M params
Tasks /extract
Licenseapache-2.0
Languagesen
Latency1.6 s
Throughput510 tok/s
Cost$0.436 /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Extraction

Output kindstext_regions
Inputstext · image
Max sequence length

Benchmarks

DocVQA

general kie en

Visual question answering on document images

Corpus: 5,188 Queries: 5,188
Quality
anls 0.3521
exact match 0.2600
Performance L4 b1 c16
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 2.8K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.