Why did we open-source our inference engine? Read the post

← Catalog

openai/whisper-large-v3-turbo

Open comparison →

Primitive: /extract · Extract · Whisper

Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al.

Multilingual

Overview

Hardware: — drives latency, throughput & cost

Size809M params
Tasks /extract
Licensemit
Languagesen, zh, de, es, ru, ko, fr, ja, pt, tr, pl, ca, nl, ar, sv, it, id, hi, fi, vi, he, uk, el, ms, cs, ro, da, hu, ta, no, th, ur, hr, bg, lt, la, mi, ml, cy, sk, te, fa, lv, bn, sr, az, sl, kn, et, mk, br, eu, is, hy, ne, mn, bs, kk, sq, sw, gl, mr, pa, si, km, sn, yo, so, af, oc, ka, be, tg, sd, gu, am, yi, lo, uz, fo, ht, ps, tk, nn, mt, sa, lb, my, bo, tl, mg, as, tt, haw, ln, ha, ba, jw, su
Latency1.0 s
Throughput
Cost /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Extraction

Output kindsText
Inputsaudio
Max sequence length448

Benchmarks

librispeech_asr

general transcription en

Speech-to-text transcription of read English audiobooks (LibriSpeech test-clean), scored by word error rate

Corpus: 2,620 Queries: 2,620
Quality
WER (lower is better) 0.0169
CER (lower is better) 0.0054
Performance L4 b1 c4
Extract p50 1.0s
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 2.8K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.