Rumi
Text to speech
The most natural Uzbek voice synthesis. Streaming, low-latency, 10 dialect voices.
Rumi generates the most natural-sounding Uzbek speech of any TTS system. Trained on 200+ hours of human-verified conversational Uzbek audio from the Biruniy Gold dataset, every training segment carries emotion labels, sentence-complete markers, and prosody tags — so the model learns how real Uzbek speakers breathe, pause, and emphasize, not just how they pronounce individual words.
- Trained on 200+ hrs of human-verified Biruniy Gold audio
- Emotion labels + prosody tags for natural intonation
- Streaming TTS for conversational voice agents
View details