Uzbek speech dataverified. proven.

200+ hours of natural conversational Uzbek — real speech from podcasts and talk shows, not scripted reading. Every word reviewed by 2+ independent native speakers. Train ASR models to 7–8% WER. Build TTS voices with natural emotion and regional accent. Download, train, deploy — no cleaning needed.

Verified Hours

200+

Transcript Accuracy

99.2%

Unique Speakers

100+

Uzbek Dialects

10

Emotion Labels

The Uzbek Data Problem

Three reasons why Uzbek ASR & TTS still fail in production.

Crowd-Sourced = Low Quality

Common Voice has Uzbek audio — but it's read-aloud sentences by random volunteers. Not how real people talk. Models trained on it: 14–24% WER.

Scraped = Unverified

Some providers scrape 1,000+ hours of Uzbek audio. But nobody checks the transcripts. Garbage in = garbage out. 16.7% WER.

Nothing For Voice AI

There is zero verified Uzbek speech data suitable for both ASR and TTS publicly available. Until now.

We built the solution: a dataset where every word is verified by humans.

What You Get

Production-ready Uzbek speech data for ASR & TTS. Download, train, deploy. No cleaning needed.

Natural Conversation

Real speech from podcasts and talk shows — not people reading scripts into a microphone.

IAA Verified

Every low-confidence word reviewed by 2+ independent native Uzbek speakers. Conflicts resolved by senior adjudicator.

Audio Quality Scored

DNSMOS quality scoring on every segment. Only MOS ≥ 3.5 included. Background noise removed with DeepFilterNet.

Speaker Diarized

100+ unique speakers labeled. Train/Val/Test split by speaker — zero data leakage.

Word-Level Timestamps

Millisecond-accurate start/end times for every word. Confidence scores included.

10 Dialects Tagged

Toshkent, Farg'ona, Samarqand, Xorazm, Buxoro, Namangan, Andijon, Qashqadaryo, Navoiy, Surxondaryo.

Code-Switching

Natural Uzbek ↔ Russian switching included. Critical for call center applications.

Emotion Labeled

Every segment classified by emotion (neutral, happy, sad, angry, surprised) with confidence scores. Build expressive TTS voices.

TTS-Ready Flags

Each segment includes sentence_complete, is_clean_speech, and speaking_rate_wpm — filter instantly for TTS training data.

Ready to Train

HuggingFace-compatible JSONL + WAV. Works with PyTorch, Transformers, faster-whisper, VITS, and Coqui TTS out of the box.

Benchmarks

The state of Uzbek speech AI, July 2026.

Biruniy v1.0 already beats every competitor on real-world WER. Biruniy Rumi extends that lead — trained on 200+ hours of human-verified data, with a post-processing pipeline nobody else has.

BenchmarksKotib AIGigaAM LargeMuxlisa AIAisha AILYNX AIOvozifyLabsUzbekVoiceAIBiruniyBiruniy
Gold v1.0
BiruniyBiruniy
Rumi
Production
Lab WER (clean)Read speech test sets6–11%7.3% (FLEURS)²Claims <10%5–6%4–5%
Real-World WERConversational speech with noise & dialects16.7% (overall)12.7% (internal)²11.4% (calls) / 13.4% (YouTube)41.7% (Telegram)¹34.5% (Telegram)¹8–9.1%5.5–6.5%
Dialect TaggingRegions coveredPartial✅ 10✅ 12
Speaker DiarizationIndividual speaker labels✅ Included✅ Included
Human VerifiedTranscript verification method✅ IAA (2+)✅ IAA (3+)
Code-Switching (UZ/RU)Mixed-language handlingPartialPartial

¹ Aisha AI 41.7% & OvozifyLabs 34.5% — tested on real Telegram voice messages (noisy, informal). Source: OvozifyLabs HuggingFace.

² GigaAM benchmarks use clean/filtered data (FLEURS, internal). Only 6–20h of real-world test data. Uzbek is a side language, not primary focus.

Biruniy v1.0 & Rumi on held-out Biruniy Gold Test Set.
Last updated July 28, 2026.

Lab WER lies. Real-world WER is the truth.

Aisha AI hits 41.7% WER on real Telegram voice messages. OvozifyLabs: 34.5%. GigaAM scores 7.3% on FLEURS but only 12.7% on its own internal real-world test. Biruniy v1.0: 7–8% on noisy, dialectal speech — the conditions your users actually record in.

The post-processing gap

Every competitor outputs raw, unformatted text. Biruniy v1.0 (and Rumi) ship a post-processing pipeline — punctuation restoration, number formatting, spelling correction, dialect normalization — for an additional 3–5% absolute WER improvement and production-ready output. No one else has this.

Pipeline

How we build the dataset.

6 stages. 3 AI models. 2+ human reviewers per segment.

01Source

Audio from Talabam.com — Uzbekistan's largest podcast and video platform. 100% native speakers. Real conversations.

02Clean

DeepFilterNet v3 noise removal. DNSMOS quality scoring. Only MOS ≥ 3.5 passes.

03Segment

Pyannote 3.1 speaker diarization. VAD filtering (≥50% speech). 5–30 second segments with speaker labels.

04Transcribe

Whisper ASR with word-level timestamps. Confidence scoring on every word. Low-confidence words auto-flagged for review.

05Human Verify

Every flagged word reviewed by 2+ independent native Uzbek speakers (Inter-Annotator Agreement). Conflicts resolved by senior adjudicator. This is what makes Biruniy data gold-standard.

06Export

HuggingFace-compatible JSONL + WAV. Speaker-based train/val/test splits. Full metadata: dialect, MOS, speaker ID, timestamps.

Use Cases

Who uses Biruniy data.

Banking & Finance

Build voice agents for Uzbek banking customers. Understand real conversational Uzbek — not scripted prompts. Code-switching support for Uzbek-Russian bilingual users.

Call Centers

Automate call transcription and quality monitoring. 10 dialect regions means your model works across all of Uzbekistan, not just Tashkent.

AI Labs & Researchers

Train or fine-tune Uzbek ASR and TTS models. Production-ready format. Zero preprocessing needed. Speaker-based splits prevent data leakage.

Voice Assistants & TTS

Build Uzbek voice assistants that sound natural. Our speaker-isolated, emotion-labeled data trains TTS models that speak with the right tone, pace, and regional accent.

Voice Cloning & Dubbing

100+ unique speakers with consistent audio quality. Perfect for voice cloning, audiobook narration, and video dubbing in Uzbek.

Dataset Access

Get the data that powers 7–8% WER.

Contact for pricing

Research License

  • 200+ hours verified conversational Uzbek
  • HuggingFace format (JSONL + WAV)
  • Speaker-based train/val/test splits
  • Full metadata (dialect, MOS, timestamps)
  • TTS quality flags (sentence_complete, emotion, clean_speech)
  • Topic classification per segment
  • Commercial use allowed
  • Email support
Custom contract

Enterprise

  • Everything in Research License
  • Custom data collection (your domain vocabulary)
  • Scale to 200+ hours on demand
  • Domain-specific: banking, medical, legal, telecom
  • Custom TTS voice data collection
  • Single-speaker datasets on demand
  • On-premise delivery option
  • Dedicated account manager
  • SLA guarantee

Contact

Get the data that powers 7–8% WER and natural-sounding Uzbek TTS.

Tell us what you're building. We'll respond within 24 hours.