Uzbek STT
Nava: Uzbek Speech-to-Text with 7–8% Real-World WER
Nava delivers industry-leading Uzbek transcription accuracy across real-world call center, media, and conversational voice agent audio.
Today we are introducing Nava, Biruniy's flagship Uzbek speech-to-text model. Nava is designed for production applications that require accurate transcription of natural, conversational Uzbek speech across real-world acoustic environments.
While previous academic models or scraped datasets struggle with conversational audio (exhibiting 16.7% to 25%+ Word Error Rates), Nava achieves 7–8% WER on conversational Uzbek and 5–6% on clean read speech.
Why existing Uzbek speech models fail in production
Most speech models tested on benchmark datasets fall apart when exposed to actual phone calls, podcasts, or customer service interactions in Uzbekistan. Three primary factors cause this breakdown:
- Scraped and unverified data: Models trained on thousands of hours of unverified web scrapes inherit misspellings, hallucinations, and inaccurate labels.
- Monodialect assumptions: Models trained exclusively on standard literary Tashkent Uzbek fail when confronted with speakers from Samarkand, Fergana, Khorezm, or Kashkadarya.
- Conversational code-switching: Everyday conversations in Uzbekistan naturally blend Uzbek and Russian vocabulary. Traditional models treat Russian loanwords or phrases as transcription errors.
Nava addresses these challenges by training on 200+ hours of gold-standard conversational audio verified by 2+ independent native speakers (Inter-Annotator Agreement).
Streaming ASR with sub-300ms latency
In conversational voice applications, latency is critical. If transcription takes more than a few hundred milliseconds after the user stops speaking, the interaction feels unnatural and disjointed.
Nava is architected for real-time streaming inference:
- Semantic endpointing: Distinguishes between brief mid-sentence pauses and actual conversational turn completions.
- Streaming partials: Emits real-time tokens as the speaker talks so downstream LLMs can pre-process context.
- Optimized ONNX runtime: Achieves first-token transcription latency under 300 ms on standard GPU hardware and efficient CPU fallback.
Comprehensive regional coverage
Nava is trained on speech across all 10 major dialect regions of Uzbekistan:
- Tashkent City & Region
- Fergana Valley (Fergana, Andijan, Namangan)
- Samarkand & Bukhara
- Khorezm
- Southern regions (Kashkadarya, Surkhandarya, Navoi)
The model natively recognizes dialect-specific phonetic shifts and lexical variants without requiring separate models for each region.
Nava model overview
| Capability | Details |
|---|---|
| Modality | Speech → Text (ASR) |
| Real-world WER | 7–8% on conversational Uzbek |
| Clean WER | 5–6% on studio/clean audio |
| Dialects | 10 Uzbek dialect regions |
| Code-switching | Native Uzbek ↔ Russian |
| Latency | < 300 ms streaming latency |
| Runtimes | Transformers, faster-whisper, ONNX Runtime |
| Access | Biruniy Studio API & On-Premise Enterprise |
Try Nava in Biruniy Studio
You can test Nava directly in Biruniy Studio or test the microphone demo on our STT product page.
For enterprise on-premise deployments, VPC installations, or domain-specific fine-tuning on your company's audio, contact our sales team.
Related content
Meeting Intelligence
Uzbek Meeting Intelligence: Speaker Diarization and AI Summaries
Transform Uzbek business meetings, calls, and interviews into structured transcripts with millisecond-accurate speaker diarization and automated executive summaries.
Uzbek TTS
Rumi: Natural Uzbek Text-to-Speech
Rumi turns written Uzbek into natural speech with streaming inference, regional voice coverage, and emotion-aware prosody for production voice applications.