Kokoro TTS
Model, speech, by hexgrad
Assessment
The best-sounding voice for its size by a distance, and the one voice engine that needs no GPU mode switch because it fits beside a loaded LLM. The right pick for narration and Studio work; Piper remains the one for calls.
2026-10-09
Strengths
- Natural, expressive English voices for a model of about 82M parameters
- Light enough to sit beside a loaded language model on the same card
- Apache-licensed weights with a simple Python package
Limitations
- A handful of voices, mostly English
- Needs a GPU for real-time; on a CPU it is usable but slow for long text
Kokoro is a small open text-to-speech model, about 82 million parameters, whose voices sound far more natural than the size suggests.[1] It ships as a Python package with a few dozen voices, mostly English, and permissive weights.
What it is for
Narration, voice-over scripts, anything spoken for people to listen to rather than merely hear. Where Piper is the reliable call voice, Kokoro is the pleasant one.
Sharing the card
Measured beside a loaded 26B language model on a 12 GB card, Kokoro added about 1.2 GB of VRAM and no measurable slowdown to the model's decoding.[2] That makes it the one voice engine that can run without taking the cards away from the LLM; heavier engines such as Seed-VC or Chatterbox either slow the model sharply or do not fit at all and need a single-card mode.
Things to know
- Pick the voice per use and keep it; mixing voices in one piece reads as a mistake.
- Chunk long text by sentence and concatenate; it keeps timing even and memory flat.
- For a custom voice the answer is still Piper fine-tuning or a voice-conversion model, not Kokoro.
History
- 2026-10-09: article written and published.
- 2026-10-09: entry created.
Comments
Public comments on each entry are coming. Nothing is collected here yet.