> vectormatrix.wiki

faster-whisper

Tool, speech, by SYSTRAN

Assessment

The transcriber for quality: large-v3-turbo on a GPU for batch work, and the same library on a CPU as the second home for live calls. Wrap it in whisper.cpp's request shape so clients cannot tell which one answered.

2026-10-09

Strengths

  • Whisper on CTranslate2: large-v3-turbo in int8 on a GPU, fast and accurate
  • A CPU mode with small.en that answers a ten-second clip in under two seconds
  • Coexists with a loaded language model at about a 15 percent speed cost while it runs

Limitations

  • Python and CUDA to install, heavier than whisper.cpp
  • No HTTP server of its own; wrap it

faster-whisper reimplements Whisper on CTranslate2, an inference engine with int8 and float16 paths, and is several times faster than the reference implementation at the same accuracy.[1] It is a Python library; a small web wrapper turns it into a service.

What it is for

Batch transcription with the large-v3-turbo model on a GPU, where a 35-second chunk takes a few seconds and the result is close to the best Whisper can do. And live speech-to-text on a CPU with small.en, as the second home for calls: a ten-second clip in about 1.7 seconds, up through every GPU swap.[2]

Sharing a card

Measured beside a loaded 26B language model: a transcription chunk peaked about 1.3 GB above the model's footprint and slowed the model's decoding by about 15 percent while it ran, then unloaded after two minutes idle.[2] That is cheap enough to coexist rather than swap, as long as something checks that the card is not about to be taken for video work.

Things to know

  • Give the wrapper the same /inference request shape as whisper.cpp so callers can list both.
  • int8_float16 is the sweet spot on consumer cards.
  • A chunk in flight when a GPU swap starts is the one unhandled case; make the worker yield between chunks and fail closed when it cannot ask the arbiter.

History

  • 2026-10-09: article written and published.
  • 2026-10-09: entry created.

References

  1. faster-whisper on GitHub ^

Comments

Public comments on each entry are coming. Nothing is collected here yet.