Nava: Speech to text for Uzbek.

Ranked #1 on Uzbek conversational accuracy, built for voice agents with semantic endpointing and industry-leading latency. Achieves 8–9.1% WER on real-world Uzbek audio across 10 dialects.

Language: Uzbek

Click the microphone to start transcribing

Click or press Space to start a transcription

Use Cases

One transcription model for every
environment your business takes you to.

Call Centers

Transcribe Uzbek customer calls in real time. Handle telephony audio, background noise, and natural conversation flow — all with 8–9.1% WER accuracy.

Voice Agents

Power Uzbek voice assistants with streaming ASR. Partial results, turn detection, and code-switching support make conversational AI feel natural in Uzbek.

Media & Podcasts

Transcribe Uzbek podcasts, interviews, and broadcasts. Speaker diarization and dialect tagging give you structured, searchable archives of Uzbek audio content.

Banking & Finance

Handle domain-specific Uzbek vocabulary — financial terms, account numbers, and currency amounts. Accurate transcription for compliance, analytics, and customer experience.

Capabilities

Built for Uzbek Voice AI.

Four capabilities that make Nava the transcription layer production Uzbek voice agents rely on.

Accuracy

Heard right the first time.

In practice

In production, the transcript is the foundation everything else builds on. A transcription error undermines the LLM input and takes the interaction in the wrong direction. The inverse is equally true — accuracy compounds, and a precise transcript means a better response and a call that resolves.

Nava's approach

Nava achieves 8–9.1% Word Error Rate on conversational Uzbek — the lowest of any Uzbek ASR model. Trained on human-verified data with IAA quality assurance, it natively handles code-switching between Uzbek and Russian, phone numbers, dates, and named entities. Built for real-world audio — telephony, background noise, and all 10 major Uzbek dialects.

Conversational Flow

Knows when you start and finish.

In practice

A conversation has two critical moments — when a caller starts talking and when they finish. Miss the start and the agent misses the turn entirely. Trigger too early on the end and the agent jumps in mid-thought. The right transcription model gets both right without the wait.

Nava's approach

Nava supports streaming transcription with real-time turn detection. Built-in endpointing determines turn end by meaning, not silence — so natural pauses mid-thought don't trigger the agent prematurely. Partial results stream as the speaker talks, keeping latency minimal from last word to first response.

Speed

The caller stops talking. The agent starts thinking.

In practice

When transcription is fast and consistent, the agent's response feels immediate. One slow transcript in ten means one call in ten where that readiness breaks. Nine great calls don't cancel out the one that didn't feel right.

Nava's approach

Nava is optimized for fast streaming inference. Combined with ONNX runtime optimization, streaming latency is under 0.3 seconds. Partial results arrive even faster, so your voice agent responds without perceptible lag.

Cost

Quality that scales with you.

In practice

Voice is the most natural interface for communication. Getting cost and quality right at scale enables voice everywhere — the default interface across every agentic interaction. You shouldn't have to choose between accuracy and affordability.

Nava's approach

Nava is released under Apache 2.0 — no per-call pricing, no API tokens, no usage caps. Run it on your own infrastructure at your own scale. Efficient inference keeps costs low at high volume. Fine-tune further on your domain data without starting from scratch.

Performance

Fast. Open.
Apache 2.0.

Nava is built for production: streaming partial results, sub-300ms latency, and open weights. No API keys. No usage limits.

Streaming latency

< 0.3 s

Fine-tuning

Supported

License

Apache 2.0

Inference

Transformers / ONNX

Input

16kHz mono WAV

Output

Text + timestamps

FAQ

Frequently asked questions.

What is Nava?

Nava is a fine-tuned Uzbek speech-to-text model. It achieves 8–9.1% Word Error Rate (WER) on conversational Uzbek speech — the lowest of any publicly available Uzbek ASR model. It supports streaming transcription, 10 Uzbek dialects, and code-switching between Uzbek and Russian.

How does Nava compare to other Uzbek ASR models?

Nava achieves 8–9.1% WER on conversational Uzbek. By comparison, models trained on unverified scraped data achieve 16.7% WER, and models trained on Common Voice Uzbek achieve 14-24% WER. The key difference is training data quality — Biruniy is trained on human-verified data with IAA quality assurance.

Does Nava support streaming (real-time) transcription?

Yes. Nava supports streaming ASR with partial results, turn detection, and endpointing. Latency from end of speech to final transcript is under 0.3 seconds with ONNX runtime optimization. This makes it suitable for real-time voice agents, call center transcription, and live captioning.

Which Uzbek dialects does Nava support?

Nava is trained on data tagged across 10 major Uzbek dialect regions: Toshkent, Farg'ona, Samarqand, Xorazm, Buxoro, Namangan, Andijon, Qashqadaryo, Navoiy, and Surxondaryo. The model handles dialectal variation in pronunciation and vocabulary.

Can Nava handle code-switching between Uzbek and Russian?

Yes. The training data includes natural code-switched speech (Uzbek-Russian), which is common in everyday conversation across Uzbekistan. The model transcribes code-switched audio accurately without treating Russian segments as errors.

What license is Nava released under?

Nava is released under the Apache 2.0 license. You can use it commercially, modify it, and distribute it. Run it on your own infrastructure — no per-call pricing, no API tokens, no usage caps.

Can I fine-tune Nava on my own domain data?

Yes. Nava can be fine-tuned on domain-specific data (banking, medical, legal, telecom) without retraining from scratch. Fine-tuning is included in annual and enterprise license agreements.

How do I get started with Nava?

Nava's open-source release (etamin/biruniy-v1, Apache 2.0) is coming soon — join the waitlist on the homepage to be notified. For production use today, request a license or a 30-day evaluation.

License Nava
for your use case.

We sell Nava as a model license — buy it once and run it on your own infrastructure. Use it for call centers, voice agents, transcription pipelines, banking IVR, or any production workload that needs accurate Uzbek speech recognition. No per-call fees, no API keys, no usage caps.