🛰️ Daily AI Frontier
‹ back to 2026-09-25

YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech

Research Multimodal & Generative

Ranking

Overall 78
Content 100
Popularity 27

Observed public metrics from 1 member.

Representative image for YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech

Merged summary

TL;DR - YODAS v3 is an open, weakly labeled speech corpus containing over 1.1 million hours of 48kHz multichannel audio across 147 languages. Its unprecedented scale, language coverage, and stereo fidelity could support research in multilingual speech recognition and neural audio codecs.

  • Released under CC BY 3.0 and described as the largest open speech dataset to date.
  • Introduces collection techniques designed to gather more language-balanced speech data.
  • Includes over 10,000 hours for 22 languages and over 5,000 hours for 73 languages.
  • Provides analyses of language coverage, audio quality, and transcription quality, plus baseline speech-recognition and neural-codec models.

Sources (1)

YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech

arXiv cs.CL William Chen, Shinnosuke Takamichi, Sayaka Shiota, Satoru Fukayama, Samuele Cornell, Shinji Watanabe 2026-09-24 arXiv:2609.29448
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-25 14:15:55.862825 UTC

TL;DR - YODAS v3 is an open, weakly labeled speech corpus containing over 1.1 million hours of 48kHz multichannel audio across 147 languages. Its unprecedented scale, language coverage, and stereo fidelity could support research in multilingual speech recognition and neural audio codecs.

  • Released under CC BY 3.0 and described as the largest open speech dataset to date.
  • Introduces collection techniques designed to gather more language-balanced speech data.
  • Includes over 10,000 hours for 22 languages and over 5,000 hours for 73 languages.
  • Provides analyses of language coverage, audio quality, and transcription quality, plus baseline speech-recognition and neural-codec models.
item →