openai/whisper-large-v3-turbo
Primitive: /extract · Extract ·
Whisper
Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al.
View on Hugging Face → Fine-tuned from openai/whisper-large-v3
Overview
Hardware: — drives latency, throughput & cost
| Size | 809M params |
|---|---|
| Tasks | /extract |
| License | mit |
| Languages | en, zh, de, es, ru, ko, fr, ja, pt, tr, pl, ca, nl, ar, sv, it, id, hi, fi, vi, he, uk, el, ms, cs, ro, da, hu, ta, no, th, ur, hr, bg, lt, la, mi, ml, cy, sk, te, fa, lv, bn, sr, az, sl, kn, et, mk, br, eu, is, hy, ne, mn, bs, kk, sq, sw, gl, mr, pa, si, km, sn, yo, so, af, oc, ka, be, tg, sd, gu, am, yi, lo, uz, fo, ht, ps, tk, nn, mt, sa, lb, my, bo, tl, mg, as, tt, haw, ln, ha, ba, jw, su |
| Latency | 1.0 s |
| Throughput | — |
| Cost | — /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Extraction
| Output kinds | Text |
|---|---|
| Inputs | audio |
| Max sequence length | 448 |
Benchmarks
librispeech_asr
Speech-to-text transcription of read English audiobooks (LibriSpeech test-clean), scored by word error rate