Nava: Speech to text for Uzbek.
Ranked #1 on Uzbek conversational accuracy, built for voice agents with semantic endpointing and industry-leading latency. Achieves 8–9.1% WER on real-world Uzbek audio across 10 dialects.
Language: Uzbek
Click the microphone to start transcribing
Click or press Space to start a transcription
Use Cases
One transcription model for every
environment your business takes you to.
Call Centers
Transcribe Uzbek customer calls in real time. Handle telephony audio, background noise, and natural conversation flow — all with 8–9.1% WER accuracy.
Voice Agents
Power Uzbek voice assistants with streaming ASR. Partial results, turn detection, and code-switching support make conversational AI feel natural in Uzbek.
Media & Podcasts
Transcribe Uzbek podcasts, interviews, and broadcasts. Speaker diarization and dialect tagging give you structured, searchable archives of Uzbek audio content.
Banking & Finance
Handle domain-specific Uzbek vocabulary — financial terms, account numbers, and currency amounts. Accurate transcription for compliance, analytics, and customer experience.
Capabilities
Built for Uzbek Voice AI.
Four capabilities that make Nava the transcription layer production Uzbek voice agents rely on.
Accuracy
Heard right the first time.
In practice
In production, the transcript is the foundation everything else builds on. A transcription error undermines the LLM input and takes the interaction in the wrong direction. The inverse is equally true — accuracy compounds, and a precise transcript means a better response and a call that resolves.
Nava's approach
Nava achieves 8–9.1% Word Error Rate on conversational Uzbek — the lowest of any Uzbek ASR model. Trained on human-verified data with IAA quality assurance, it natively handles code-switching between Uzbek and Russian, phone numbers, dates, and named entities. Built for real-world audio — telephony, background noise, and all 10 major Uzbek dialects.
Conversational Flow
Knows when you start and finish.
In practice
A conversation has two critical moments — when a caller starts talking and when they finish. Miss the start and the agent misses the turn entirely. Trigger too early on the end and the agent jumps in mid-thought. The right transcription model gets both right without the wait.
Nava's approach
Nava supports streaming transcription with real-time turn detection. Built-in endpointing determines turn end by meaning, not silence — so natural pauses mid-thought don't trigger the agent prematurely. Partial results stream as the speaker talks, keeping latency minimal from last word to first response.
Speed
The caller stops talking. The agent starts thinking.
In practice
When transcription is fast and consistent, the agent's response feels immediate. One slow transcript in ten means one call in ten where that readiness breaks. Nine great calls don't cancel out the one that didn't feel right.
Nava's approach
Nava is optimized for fast streaming inference. Combined with ONNX runtime optimization, streaming latency is under 0.3 seconds. Partial results arrive even faster, so your voice agent responds without perceptible lag.
Cost
Quality that scales with you.
In practice
Voice is the most natural interface for communication. Getting cost and quality right at scale enables voice everywhere — the default interface across every agentic interaction. You shouldn't have to choose between accuracy and affordability.
Nava's approach
Nava is released under Apache 2.0 — no per-call pricing, no API tokens, no usage caps. Run it on your own infrastructure at your own scale. Efficient inference keeps costs low at high volume. Fine-tune further on your domain data without starting from scratch.
Performance
Fast. Open.
Apache 2.0.
Nava is built for production: streaming partial results, sub-300ms latency, and open weights. No API keys. No usage limits.
Streaming latency
< 0.3 s
Fine-tuning
Supported
License
Apache 2.0
Inference
Transformers / ONNX
Input
16kHz mono WAV
Output
Text + timestamps
FAQ
Frequently asked questions.
What is Nava?
Nava is a fine-tuned Uzbek speech-to-text model. It achieves 8–9.1% Word Error Rate (WER) on conversational Uzbek speech — the lowest of any publicly available Uzbek ASR model. It supports streaming transcription, 10 Uzbek dialects, and code-switching between Uzbek and Russian.
How does Nava compare to other Uzbek ASR models?
Nava achieves 8–9.1% WER on conversational Uzbek. By comparison, models trained on unverified scraped data achieve 16.7% WER, and models trained on Common Voice Uzbek achieve 14-24% WER. The key difference is training data quality — Biruniy is trained on human-verified data with IAA quality assurance.
Does Nava support streaming (real-time) transcription?
Yes. Nava supports streaming ASR with partial results, turn detection, and endpointing. Latency from end of speech to final transcript is under 0.3 seconds with ONNX runtime optimization. This makes it suitable for real-time voice agents, call center transcription, and live captioning.
Which Uzbek dialects does Nava support?
Nava is trained on data tagged across 10 major Uzbek dialect regions: Toshkent, Farg'ona, Samarqand, Xorazm, Buxoro, Namangan, Andijon, Qashqadaryo, Navoiy, and Surxondaryo. The model handles dialectal variation in pronunciation and vocabulary.
Can Nava handle code-switching between Uzbek and Russian?
Yes. The training data includes natural code-switched speech (Uzbek-Russian), which is common in everyday conversation across Uzbekistan. The model transcribes code-switched audio accurately without treating Russian segments as errors.
What license is Nava released under?
Nava is released under the Apache 2.0 license. You can use it commercially, modify it, and distribute it. Run it on your own infrastructure — no per-call pricing, no API tokens, no usage caps.
Can I fine-tune Nava on my own domain data?
Yes. Nava can be fine-tuned on domain-specific data (banking, medical, legal, telecom) without retraining from scratch. Fine-tuning is included in annual and enterprise license agreements.
How do I get started with Nava?
Nava's open-source release (etamin/biruniy-v1, Apache 2.0) is coming soon — join the waitlist on the homepage to be notified. For production use today, request a license or a 30-day evaluation.
License Nava
for your use case.
We sell Nava as a model license — buy it once and run it on your own infrastructure. Use it for call centers, voice agents, transcription pipelines, banking IVR, or any production workload that needs accurate Uzbek speech recognition. No per-call fees, no API keys, no usage caps.