whisper.cpp
Tool, speech, by ggerganov
Assessment
The dependable live transcriber: the same request shape on a 4 GB AMD card and on a CPU, so a caller can list two homes and fail over. Use the small English model for calls and a larger one in batch where time does not matter.
2026-10-09
Strengths
- Live speech-to-text on small or odd GPUs through Vulkan, and on CPUs
- One binary, one HTTP endpoint, the same request shape everywhere
- The small English model answers a ten-second clip in a couple of seconds on modest hardware
Limitations
- Accuracy on crosstalk and accents trails the large models
- A GPU build shares the card with everything else and slows when the card fills
whisper.cpp is the C/C++ port of OpenAI's Whisper speech-recognition models, from the same author as llama.cpp.[1] It runs on CPUs and, through Vulkan, Metal or CUDA, on GPUs that the Python stack cannot use, including small AMD cards. Its server example answers a multipart POST of audio with JSON text.
What it is for
Live calls: a caller speaks, a clip is posted, text comes back in a second or two, and the conversation goes on. Batch transcription of recordings with a larger model when accuracy matters more than latency.
Two homes, one shape
Because the request is the same on every build, a client can keep a list of servers and try the next when one does not answer.[2] A GPU build on one machine and a CPU build on another is the pattern: the call is always heard, even while the GPU box is busy or down.
Things to know
- On a shared small card, a full card slows transcription about twenty times while every health check says fine; health-check with a real clip, not a ping.
small.enis the right model for live English;large-v3-turbofor batch, in faster-whisper or here.- Silence detection in front of the model saves more time than any other tuning.
History
- 2026-10-09: article written and published.
- 2026-10-09: entry created.
Comments
Public comments on each entry are coming. Nothing is collected here yet.