The lineup

Two production-ready Uzbek speech AI models.

Rumi synthesizes the most natural Uzbek speech. Nava transcribes the most accurate conversational Uzbek. Both built on the same verified Biruniy Gold dataset.

Available models

Speech synthesis and transcription for production.

Choose Rumi for natural Uzbek voice synthesis or Nava for accurate conversational transcription. Open each model to compare its capabilities, deployment profile, and best-fit use cases.

Rumi

Text to speech

Production

The most natural Uzbek voice synthesis. Streaming, low-latency, 10 dialect voices.

Rumi generates the most natural-sounding Uzbek speech of any TTS system. Trained on 200+ hours of human-verified conversational Uzbek audio from the Biruniy Gold dataset, every training segment carries emotion labels, sentence-complete markers, and prosody tags — so the model learns how real Uzbek speakers breathe, pause, and emphasize, not just how they pronounce individual words.

  • Trained on 200+ hrs of human-verified Biruniy Gold audio
  • Emotion labels + prosody tags for natural intonation
  • Streaming TTS for conversational voice agents

View details

License any model
for your use case.

Buy once, run on your own infrastructure. No per-call fees, no API keys, no usage caps.