faster-whisper
Tool, speech, by SYSTRAN
Assessment
The transcriber for quality: large-v3-turbo on a GPU for batch work, and the same library on a CPU as the second home for live calls. Wrap it in whisper.cpp's request shape so clients cannot tell which one answered.
2026-10-09
Strengths
- Whisper on CTranslate2: large-v3-turbo in int8 on a GPU, fast and accurate
- A CPU mode with small.en that answers a ten-second clip in under two seconds
- Coexists with a loaded language model at about a 15 percent speed cost while it runs
Limitations
- Python and CUDA to install, heavier than whisper.cpp
- No HTTP server of its own; wrap it
faster-whisper reimplements Whisper on CTranslate2, an inference engine with int8 and float16 paths, and is several times faster than the reference implementation at the same accuracy.[1] It is a Python library; a small web wrapper turns it into a service.
What it is for
Batch transcription with the large-v3-turbo model on a GPU, where a 35-second chunk takes a few seconds and the result is close to the best Whisper can do. And live speech-to-text on a CPU with small.en, as the second home for calls: a ten-second clip in about 1.7 seconds, up through every GPU swap.[2]
Sharing a card
Measured beside a loaded 26B language model: a transcription chunk peaked about 1.3 GB above the model's footprint and slowed the model's decoding by about 15 percent while it ran, then unloaded after two minutes idle.[2] That is cheap enough to coexist rather than swap, as long as something checks that the card is not about to be taken for video work.
Things to know
- Give the wrapper the same
/inferencerequest shape as whisper.cpp so callers can list both. - int8_float16 is the sweet spot on consumer cards.
- A chunk in flight when a GPU swap starts is the one unhandled case; make the worker yield between chunks and fail closed when it cannot ask the arbiter.
History
- 2026-10-09: article written and published.
- 2026-10-09: entry created.
Comments
Public comments on each entry are coming. Nothing is collected here yet.