> vectormatrix.wiki

Kokoro TTS

Model, speech, by hexgrad

Assessment

The best-sounding voice for its size by a distance, and the one voice engine that needs no GPU mode switch because it fits beside a loaded LLM. The right pick for narration and Studio work; Piper remains the one for calls.

2026-10-09

Strengths

  • Natural, expressive English voices for a model of about 82M parameters
  • Light enough to sit beside a loaded language model on the same card
  • Apache-licensed weights with a simple Python package

Limitations

  • A handful of voices, mostly English
  • Needs a GPU for real-time; on a CPU it is usable but slow for long text

Kokoro is a small open text-to-speech model, about 82 million parameters, whose voices sound far more natural than the size suggests.[1] It ships as a Python package with a few dozen voices, mostly English, and permissive weights.

What it is for

Narration, voice-over scripts, anything spoken for people to listen to rather than merely hear. Where Piper is the reliable call voice, Kokoro is the pleasant one.

Sharing the card

Measured beside a loaded 26B language model on a 12 GB card, Kokoro added about 1.2 GB of VRAM and no measurable slowdown to the model's decoding.[2] That makes it the one voice engine that can run without taking the cards away from the LLM; heavier engines such as Seed-VC or Chatterbox either slow the model sharply or do not fit at all and need a single-card mode.

Things to know

  • Pick the voice per use and keep it; mixing voices in one piece reads as a mistake.
  • Chunk long text by sentence and concatenate; it keeps timing even and memory flat.
  • For a custom voice the answer is still Piper fine-tuning or a voice-conversion model, not Kokoro.

History

  • 2026-10-09: article written and published.
  • 2026-10-09: entry created.

References

  1. Kokoro on GitHub ^
  2. Kokoro on Hugging Face ^

Comments

Public comments on each entry are coming. Nothing is collected here yet.