Rumi: The most natural Uzbek text to speech model.

Ranked #1 for naturalness, sub-300ms first-byte latency, and natively Uzbek across 10 regional dialects — built for production voice agents.

115/500

Use Cases

One TTS model for every
environment your business takes you to.

Voice Agents

Stream low-latency Uzbek speech to power conversational AI assistants. Natural intonation, code-switching support, and fast first-byte audio make every turn feel human.

Audiobooks & Media

Convert Uzbek text into natural narration for audiobooks, news, e-learning, and broadcast content. Multiple regional voices let you match content to the audience.

Banking & Finance

Generate clear, professional Uzbek voice prompts for IVR systems, account notifications, and transaction confirmations. Reliable pronunciation of numbers, names, and amounts.

Accessibility

Read Uzbek text aloud for visually impaired users, language learners, and anyone who prefers listening. Natural prosody makes long-form content comfortable to follow.

Capabilities

Built for Uzbek Voice AI.

Four capabilities that make Rumi the speech synthesis layer production Uzbek voice products rely on.

Naturalness

Sounds like an Uzbek speaker, not a translator.

In practice

The moment a synthetic voice sounds synthetic, trust breaks. A robotic accent on an Uzbek word pulls the listener out of the conversation and reminds them they're talking to software. Pronunciation, stress, and intonation need to feel native — or the rest of the system never gets a fair hearing.

Rumi's approach

Rumi is trained on 200+ hours of human-verified conversational Uzbek audio from the Biruniy Gold dataset. Every segment carries emotion labels, sentence-complete markers, and prosody tags — so the model learns how real Uzbek speakers breathe, pause, and emphasize, not just how they pronounce individual words.

Streaming

First word plays before the last word is generated.

In practice

Voice agents feel natural when the first word reaches the caller fast. If the user has to wait until the entire sentence is rendered before audio begins, every turn carries a noticeable delay. Streaming TTS is the difference between a conversation and a Q&A form.

Rumi's approach

Rumi supports streaming inference — audio chunks are produced as the model decodes, so the first syllable plays while the rest of the sentence is still being generated. End-to-end latency from text input to first audio is under 300ms, so voice agents respond without perceptible lag.

Dialect Coverage

Native voices for every region of Uzbekistan.

In practice

Uzbek has real dialectal variation. A model trained on one region's intonation reads as foreign to speakers from another. Banking customers in Xorazm, callers from Farg'ona, government communications in Toshkent — each audience expects to hear a voice that matches their context.

Rumi's approach

Rumi ships with 5+ voice profiles covering the major regional intonations of Uzbek: Toshkent, Farg'ona, Samarqand, and more. Each voice is trained on dialect-tagged data from the Biruniy Gold dataset, so pronunciation and prosody match the regional speech patterns native listeners expect.

Cost

Quality voice synthesis that scales with you.

In practice

Voice is the most natural interface for communication. Getting cost and quality right at scale enables voice everywhere — the default interface across every agentic interaction. You shouldn't have to choose between naturalness and affordability.

Rumi's approach

Rumi is released as a model license — buy it once and run it on your own infrastructure. No per-character pricing, no API tokens, no usage caps. Self-host it, fine-tune it, deploy it across thousands of concurrent streams — voice synthesis that fits the economics of production.

Architecture

Neural TTS. Apache 2.0.

A neural text-to-speech architecture trained on human-verified conversational Uzbek audio. Open weights. No API keys. No usage limits. Stream low-latency speech for any production workload.

Base Model

Neural TTS + HiFi-GAN

Training Data

Biruniy Gold (200+ hrs)

License

Apache 2.0

Inference

Streaming / Batch

First-byte Latency

< 300 ms

Output

WAV / MP3 24kHz

FAQ

Frequently asked questions.

What is Rumi?

Rumi is a fine-tuned Uzbek text-to-speech model built for natural voice synthesis. It produces the most natural-sounding Uzbek speech of any TTS system, with support for 10 regional dialect voices, streaming inference, and code-switching between Uzbek and Russian.

How natural does Rumi sound compared to other TTS systems?

Rumi is trained on 200+ hours of human-verified conversational Uzbek audio from the Biruniy Gold dataset. Every training segment carries emotion labels, sentence-complete markers, and prosody tags — so the model learns how real Uzbek speakers talk, not just how they pronounce individual words. Subjective listening tests rate Rumi as significantly more natural than general-purpose multilingual TTS systems.

Does Rumi support streaming (real-time) TTS?

Yes. Rumi supports streaming inference — audio chunks are produced as the model decodes, so the first syllable plays while the rest of the sentence is still being generated. End-to-end latency from text input to first audio is under 300ms with optimized inference runtimes. This makes it suitable for real-time voice agents and conversational AI.

Which Uzbek dialects and voices does Rumi support?

Rumi ships with 5+ voice profiles covering the major regional intonations of Uzbek: Toshkent, Farg'ona, Samarqand, and more. Each voice is trained on dialect-tagged data, so pronunciation and prosody match the regional speech patterns native listeners expect.

Can Rumi handle code-switching between Uzbek and Russian?

Yes. The training data includes natural code-switched text (Uzbek-Russian), which is common in everyday communication across Uzbekistan. Rumi pronounces Russian words and phrases embedded in Uzbek sentences with native-like fluency rather than treating them as errors.

How is Rumi licensed?

Rumi is sold as a model license. You buy it once and run it on your own infrastructure for your use case — voice agents, IVR, audiobooks, accessibility, or any production workload where natural Uzbek speech matters. No per-character fees, no API keys, no usage caps.

Can I fine-tune Rumi on my own voice data?

Yes. Rumi's architecture supports fine-tuning on domain-specific voice data (brand voices, celebrity voices, custom accents) without retraining from scratch. Provide 30+ minutes of clean, transcribed audio in your target voice and Rumi adapts.

How do I get started with Rumi?

Request a license and we'll get you up and running — model weights, integration guide, voice configuration, and reference code for streaming inference. Visit the contact page to start a conversation about your use case.

License Rumi
for your use case.

We sell Rumi as a model license — buy it once and run it on your own infrastructure. Use it for voice agents, IVR, audiobooks, accessibility, or any production workload where natural Uzbek speech matters. No per-character fees, no API keys, no usage caps.