microsoft/Florence-2-base-ft
Primitive: /extract · Extract ·
Florence-2
This Hub repository contains a HuggingFace's `transformers` implementation of Florence-2 model from Microsoft.
MultimodalText regions
Overview
Hardware: — drives latency, throughput & cost
| Size | 232M params |
|---|---|
| Tasks | /extract |
| License | mit |
| Latency | — |
| Throughput | — |
| Cost | — /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Extraction
| Output kinds | text_regions |
|---|---|
| Inputs | image |
| Max sequence length | — |
Benchmarks
COCO
Object detection on COCO natural images
Corpus: 5,000 Queries: 5,000
Quality
ap 0.2947
ap50 0.4221
ap75 0.3144
ar 100 0.4738
Compare (0)Compare →