Why did we open-source our inference engine? Read the post

← Catalog

microsoft/Florence-2-large

Open comparison →

Primitive: /extract · Extract · Florence-2

This is a continued pretrained version of Florence-2-large model with 4k context length, only 0.1B samples are used for continue pretraining, thus it might not be trained well. In addition, OCR task has been updated with line separator ('\n'). COCO OD AP 39.8

MultimodalText regions

Overview

Hardware: — drives latency, throughput & cost

Size777M params
Tasks /extract
Licensemit
Latency
Throughput
Cost /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Extraction

Output kindstext_regions
Inputsimage
Max sequence length

Benchmarks

COCO

general detection en

Object detection on COCO natural images

Corpus: 5,000 Queries: 5,000
Quality
ap 0.2940
ap50 0.4216
ap75 0.3160
ar 100 0.4918
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 2.8K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.