Blog

Uzbek STT

Nava: Uzbek Speech-to-Text with 7–8% Real-World WER

Nava delivers industry-leading Uzbek transcription accuracy across real-world call center, media, and conversational voice agent audio.

·Biruniy Research

Today we are introducing Nava, Biruniy's flagship Uzbek speech-to-text model. Nava is designed for production applications that require accurate transcription of natural, conversational Uzbek speech across real-world acoustic environments.

While previous academic models or scraped datasets struggle with conversational audio (exhibiting 16.7% to 25%+ Word Error Rates), Nava achieves 7–8% WER on conversational Uzbek and 5–6% on clean read speech.

Why existing Uzbek speech models fail in production

Most speech models tested on benchmark datasets fall apart when exposed to actual phone calls, podcasts, or customer service interactions in Uzbekistan. Three primary factors cause this breakdown:

  1. Scraped and unverified data: Models trained on thousands of hours of unverified web scrapes inherit misspellings, hallucinations, and inaccurate labels.
  2. Monodialect assumptions: Models trained exclusively on standard literary Tashkent Uzbek fail when confronted with speakers from Samarkand, Fergana, Khorezm, or Kashkadarya.
  3. Conversational code-switching: Everyday conversations in Uzbekistan naturally blend Uzbek and Russian vocabulary. Traditional models treat Russian loanwords or phrases as transcription errors.

Nava addresses these challenges by training on 200+ hours of gold-standard conversational audio verified by 2+ independent native speakers (Inter-Annotator Agreement).

Streaming ASR with sub-300ms latency

In conversational voice applications, latency is critical. If transcription takes more than a few hundred milliseconds after the user stops speaking, the interaction feels unnatural and disjointed.

Nava is architected for real-time streaming inference:

  • Semantic endpointing: Distinguishes between brief mid-sentence pauses and actual conversational turn completions.
  • Streaming partials: Emits real-time tokens as the speaker talks so downstream LLMs can pre-process context.
  • Optimized ONNX runtime: Achieves first-token transcription latency under 300 ms on standard GPU hardware and efficient CPU fallback.

Comprehensive regional coverage

Nava is trained on speech across all 10 major dialect regions of Uzbekistan:

  • Tashkent City & Region
  • Fergana Valley (Fergana, Andijan, Namangan)
  • Samarkand & Bukhara
  • Khorezm
  • Southern regions (Kashkadarya, Surkhandarya, Navoi)

The model natively recognizes dialect-specific phonetic shifts and lexical variants without requiring separate models for each region.

CapabilityDetails
ModalitySpeech → Text (ASR)
Real-world WER7–8% on conversational Uzbek
Clean WER5–6% on studio/clean audio
Dialects10 Uzbek dialect regions
Code-switchingNative Uzbek ↔ Russian
Latency< 300 ms streaming latency
RuntimesTransformers, faster-whisper, ONNX Runtime
AccessBiruniy Studio API & On-Premise Enterprise

Try Nava in Biruniy Studio

You can test Nava directly in Biruniy Studio or test the microphone demo on our STT product page.

For enterprise on-premise deployments, VPC installations, or domain-specific fine-tuning on your company's audio, contact our sales team.