Piper TTS
Tool, speech, by Rhasspy
Assessment
The voice for anything that must always answer: a phone call, a reminder, a status line. It runs on any CPU, starts instantly and a custom voice is a weekend's training. Use a richer engine for Studio work where expression matters.
2026-10-09
Strengths
- Fast CPU text-to-speech: ten seconds of speech in under two seconds
- Small voice models, many languages, a one-file HTTP server
- Fine-tunable on a small corpus to make a new voice
Limitations
- Flat prosody compared with neural voices twice its size
- No streaming in the simple server; the whole WAV comes back at once
Piper is a fast, local neural text-to-speech system from the Rhasspy voice-assistant project.[1] Voices are small ONNX models, a few tens of megabytes each, and a plain HTTP server answers a POST with text by returning a 22 kHz mono WAV. On a desktop CPU it synthesizes ten seconds of speech in under two seconds.
What it is for
Speech that has to be there every time: the voice on a live phone call, a spoken reminder, a notification. Because it needs no GPU, it keeps working while the cards are busy with something else, and a second copy on another machine is cheap insurance.
Custom voices
Piper fine-tunes from a medium-quality base voice on a few hundred recorded sentences; a distinctive voice for an assistant persona is within reach of one person with a microphone and an evening.[2] The result is recognizably that person, with the engine's slightly flat delivery.
Things to know
- The server is POST-only; a bare GET answers 405, which confuses naive health checks.
- Keep sentences short; long run-on text flattens further.
- For Studio narration where expression matters, Kokoro or a heavier engine is the better tool; Piper is for reliability and speed.
History
- 2026-10-09: article written and published.
- 2026-10-09: entry created.
Comments
Public comments on each entry are coming. Nothing is collected here yet.