Twelve voices, none recorded

The model from Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech — 82M parameters, twelve voices, not one of them ever recorded. It synthesises here on two vCPU, no GPU: the same as your PC, or slower. Type Thai, pick a voice, listen.

Warming up — the model loads in the background.

What the model is actually fed

Thai is written without spaces between words, so the frontend segments it and converts to IPA before the model sees anything. If a line sounds wrong, read the phonemes first — that is usually where it went wrong.

Every voice, its own line

Each voice ships with a line written to its persona and register, and verified to phonemize cleanly before release. Pick one to load it above.