Uzbek TTS
Rumi: Natural Uzbek Text-to-Speech
Rumi turns written Uzbek into natural speech with streaming inference, regional voice coverage, and emotion-aware prosody for production voice applications.
Today we're introducing Rumi, Biruniy's Uzbek text-to-speech model for production voice applications. Rumi turns written Uzbek into natural speech with streaming inference, regional voice coverage, emotion-aware prosody, and support for Uzbek-Russian code-switching.
This post explains how Rumi is trained, why conversational data matters for Uzbek voice synthesis, and where the model fits alongside Nava, Biruniy's speech-to-text model.
Why Uzbek text-to-speech needs more than pronunciation
A voice can pronounce every word correctly and still sound wrong. Natural Uzbek speech depends on stress, rhythm, pauses, sentence endings, and the small changes in emphasis that make a voice feel present in a conversation. A model trained only on isolated or scripted sentences misses much of that context.
Rumi is built for the way people actually use voice software: talking to a voice agent, listening to an audiobook, navigating an IVR system, or hearing an accessibility feature read a long passage. The goal is not only intelligible Uzbek. It is speech that sounds native enough to keep the listener in the conversation.
Human-verified conversational training data
Rumi is trained on 50+ hours of human-verified conversational Uzbek audio from the Biruniy Gold dataset. The recordings cover natural speech rather than read-aloud prompts, with metadata that helps the model learn more than a word-to-sound mapping:
- Emotion labels help the model vary delivery with the meaning of a sentence.
- Sentence-complete segments provide clean units for training and inference.
- Prosody tags describe how speakers pause, breathe, and emphasize words.
- Dialect tags preserve regional pronunciation and intonation.
This combination gives Rumi the context needed to produce expressive Uzbek speech, not just a technically correct pronunciation of each token.
Streaming Uzbek TTS for real-time products
Rumi supports streaming inference. Audio chunks begin playing while the rest of a sentence is still being generated, instead of waiting for the complete response to finish. With optimized inference, the first audio can arrive in under 300 milliseconds.
That latency profile is designed for products where a pause feels like a broken interaction:
- Voice agents can respond without waiting for a full sentence render.
- IVR systems can start speaking while the next part of a prompt is prepared.
- Conversational AI can use natural turn-taking instead of a request-and-wait flow.
- Audiobook and accessibility applications can generate longer passages in batches.
Regional voices and Uzbek-Russian code-switching
Uzbek speech is not a single regional accent. Rumi's training data covers 10 Uzbek dialect regions, including Toshkent, Farg'ona, Samarqand, and more. The model ships with 5+ voice profiles for the major regional intonations represented in the data.
Rumi also handles natural code-switching between Uzbek and Russian. That matters for everyday speech in Uzbekistan, where borrowed words and short Russian phrases often appear inside an otherwise Uzbek sentence. The model is designed to pronounce those phrases as part of the conversation rather than treating them as transcription errors.
Rumi model profile
| Capability | Details |
|---|---|
| Model type | Neural text-to-speech with HiFi-GAN vocoder |
| Training data | Biruniy Gold, 50+ hours of verified conversational Uzbek |
| Regional coverage | 10 Uzbek dialect regions |
| Voice profiles | 5+ regional voices |
| Inference | Streaming and batch |
| First audio | Under 300 ms with optimized inference |
| Output | WAV and MP3 at 24 kHz |
| License | Apache 2.0 |
Rumi is released as a model rather than a per-character API. Teams can run it on their own infrastructure, fine-tune it for a specific domain, and scale deployment without per-call usage caps. Explore the full Rumi page for the live demo, capability details, and licensing information.
Rumi and Nava: two models for Uzbek voice AI
Rumi handles the speech synthesis side of a voice product: text in, natural Uzbek audio out. Nava handles the other half: speech in, accurate conversational Uzbek text out, with streaming partial results, semantic endpointing, and dialect-aware recognition.
Together, they cover the core speech loop for voice agents, call-center tools, banking assistants, media workflows, and accessibility products. View all Biruniy models to compare the lineup and choose the model that fits your workload.
Try Rumi in your product
You can try Rumi's Uzbek TTS demo to hear the voice and test streaming synthesis. If you are building a production voice application and need model weights, deployment guidance, or a commercial license, contact Biruniy Research.
Rumi is built around a simple idea: Uzbek voice AI should sound like it belongs in Uzbekistan. Better data, regional coverage, and low-latency inference make that standard practical for the products people use every day.