Uzbek speech dataverified. proven.
200+ hours of natural conversational Uzbek — real speech from podcasts and talk shows, not scripted reading. Every word reviewed by 2+ independent native speakers. Train ASR models to 7–8% WER. Build TTS voices with natural emotion and regional accent. Download, train, deploy — no cleaning needed.
Verified Hours
200+
Transcript Accuracy
99.2%
Unique Speakers
100+
Uzbek Dialects
10
Emotion Labels
✓
The Uzbek Data Problem
Three reasons why Uzbek ASR & TTS still fail in production.
Crowd-Sourced = Low Quality
Common Voice has Uzbek audio — but it's read-aloud sentences by random volunteers. Not how real people talk. Models trained on it: 14–24% WER.
Scraped = Unverified
Some providers scrape 1,000+ hours of Uzbek audio. But nobody checks the transcripts. Garbage in = garbage out. 16.7% WER.
Nothing For Voice AI
There is zero verified Uzbek speech data suitable for both ASR and TTS publicly available. Until now.
We built the solution: a dataset where every word is verified by humans.
What You Get
Production-ready Uzbek speech data for ASR & TTS. Download, train, deploy. No cleaning needed.
Natural Conversation
Real speech from podcasts and talk shows — not people reading scripts into a microphone.
IAA Verified
Every low-confidence word reviewed by 2+ independent native Uzbek speakers. Conflicts resolved by senior adjudicator.
Audio Quality Scored
DNSMOS quality scoring on every segment. Only MOS ≥ 3.5 included. Background noise removed with DeepFilterNet.
Speaker Diarized
100+ unique speakers labeled. Train/Val/Test split by speaker — zero data leakage.
Word-Level Timestamps
Millisecond-accurate start/end times for every word. Confidence scores included.
10 Dialects Tagged
Toshkent, Farg'ona, Samarqand, Xorazm, Buxoro, Namangan, Andijon, Qashqadaryo, Navoiy, Surxondaryo.
Code-Switching
Natural Uzbek ↔ Russian switching included. Critical for call center applications.
Emotion Labeled
Every segment classified by emotion (neutral, happy, sad, angry, surprised) with confidence scores. Build expressive TTS voices.
TTS-Ready Flags
Each segment includes sentence_complete, is_clean_speech, and speaking_rate_wpm — filter instantly for TTS training data.
Ready to Train
HuggingFace-compatible JSONL + WAV. Works with PyTorch, Transformers, faster-whisper, VITS, and Coqui TTS out of the box.
Benchmarks
The state of Uzbek speech AI, July 2026.
Biruniy v1.0 already beats every competitor on real-world WER. Biruniy Rumi extends that lead — trained on 200+ hours of human-verified data, with a post-processing pipeline nobody else has.
| Benchmarks | Kotib AI | GigaAM Large | Muxlisa AI | Aisha AI | LYNX AI | OvozifyLabs | UzbekVoiceAI | BiruniyGold v1.0 | BiruniyRumiProduction |
|---|---|---|---|---|---|---|---|---|---|
| Lab WER (clean)Read speech test sets | 6–11% | 7.3% (FLEURS)² | — | — | — | — | Claims <10% | 5–6% | 4–5% |
| Real-World WERConversational speech with noise & dialects | 16.7% (overall) | 12.7% (internal)² | 11.4% (calls) / 13.4% (YouTube) | 41.7% (Telegram)¹ | — | 34.5% (Telegram)¹ | — | 8–9.1% | 5.5–6.5% |
| Dialect TaggingRegions covered | ❌ | ❌ | Partial | ❌ | ❌ | ❌ | ❌ | ✅ 10 | ✅ 12 |
| Speaker DiarizationIndividual speaker labels | ❌ | ❌ | ❌ | ✅ | ✅ | ❌ | ✅ | ✅ Included | ✅ Included |
| Human VerifiedTranscript verification method | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ✅ IAA (2+) | ✅ IAA (3+) |
| Code-Switching (UZ/RU)Mixed-language handling | Partial | Partial | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ | ✅ |
¹ Aisha AI 41.7% & OvozifyLabs 34.5% — tested on real Telegram voice messages (noisy, informal). Source: OvozifyLabs HuggingFace.
² GigaAM benchmarks use clean/filtered data (FLEURS, internal). Only 6–20h of real-world test data. Uzbek is a side language, not primary focus.
Biruniy v1.0 & Rumi on held-out Biruniy Gold Test Set.
Last updated July 28, 2026.
Lab WER lies. Real-world WER is the truth.
Aisha AI hits 41.7% WER on real Telegram voice messages. OvozifyLabs: 34.5%. GigaAM scores 7.3% on FLEURS but only 12.7% on its own internal real-world test. Biruniy v1.0: 7–8% on noisy, dialectal speech — the conditions your users actually record in.
The post-processing gap
Every competitor outputs raw, unformatted text. Biruniy v1.0 (and Rumi) ship a post-processing pipeline — punctuation restoration, number formatting, spelling correction, dialect normalization — for an additional 3–5% absolute WER improvement and production-ready output. No one else has this.
Pipeline
How we build the dataset.
6 stages. 3 AI models. 2+ human reviewers per segment.
Audio from Talabam.com — Uzbekistan's largest podcast and video platform. 100% native speakers. Real conversations.
DeepFilterNet v3 noise removal. DNSMOS quality scoring. Only MOS ≥ 3.5 passes.
Pyannote 3.1 speaker diarization. VAD filtering (≥50% speech). 5–30 second segments with speaker labels.
Whisper ASR with word-level timestamps. Confidence scoring on every word. Low-confidence words auto-flagged for review.
Every flagged word reviewed by 2+ independent native Uzbek speakers (Inter-Annotator Agreement). Conflicts resolved by senior adjudicator. This is what makes Biruniy data gold-standard.
HuggingFace-compatible JSONL + WAV. Speaker-based train/val/test splits. Full metadata: dialect, MOS, speaker ID, timestamps.
Use Cases
Who uses Biruniy data.
Banking & Finance
Build voice agents for Uzbek banking customers. Understand real conversational Uzbek — not scripted prompts. Code-switching support for Uzbek-Russian bilingual users.
Call Centers
Automate call transcription and quality monitoring. 10 dialect regions means your model works across all of Uzbekistan, not just Tashkent.
AI Labs & Researchers
Train or fine-tune Uzbek ASR and TTS models. Production-ready format. Zero preprocessing needed. Speaker-based splits prevent data leakage.
Voice Assistants & TTS
Build Uzbek voice assistants that sound natural. Our speaker-isolated, emotion-labeled data trains TTS models that speak with the right tone, pace, and regional accent.
Voice Cloning & Dubbing
100+ unique speakers with consistent audio quality. Perfect for voice cloning, audiobook narration, and video dubbing in Uzbek.
Dataset Access
Get the data that powers 7–8% WER.
Research License
- →200+ hours verified conversational Uzbek
- →HuggingFace format (JSONL + WAV)
- →Speaker-based train/val/test splits
- →Full metadata (dialect, MOS, timestamps)
- →TTS quality flags (sentence_complete, emotion, clean_speech)
- →Topic classification per segment
- →Commercial use allowed
- →Email support
Enterprise
- →Everything in Research License
- →Custom data collection (your domain vocabulary)
- →Scale to 200+ hours on demand
- →Domain-specific: banking, medical, legal, telecom
- →Custom TTS voice data collection
- →Single-speaker datasets on demand
- →On-premise delivery option
- →Dedicated account manager
- →SLA guarantee
Contact
Get the data that powers 7–8% WER and natural-sounding Uzbek TTS.
Tell us what you're building. We'll respond within 24 hours.
