YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech
Ranking
Overall
78
Content
100
Popularity
27
Observed public metrics from 1 member.
Merged summary
TL;DR - YODAS v3 is an open, weakly labeled speech corpus containing over 1.1 million hours of 48kHz multichannel audio across 147 languages. Its unprecedented scale, language coverage, and stereo fidelity could support research in multilingual speech recognition and neural audio codecs.
- Released under CC BY 3.0 and described as the largest open speech dataset to date.
- Introduces collection techniques designed to gather more language-balanced speech data.
- Includes over 10,000 hours for 22 languages and over 5,000 hours for 73 languages.
- Provides analyses of language coverage, audio quality, and transcription quality, plus baseline speech-recognition and neural-codec models.
Sources (1)
YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - YODAS v3 is an open, weakly labeled speech corpus containing over 1.1 million hours of 48kHz multichannel audio across 147 languages. Its unprecedented scale, language coverage, and stereo fidelity could support research in multilingual speech recognition and neural audio codecs.
- Released under CC BY 3.0 and described as the largest open speech dataset to date.
- Introduces collection techniques designed to gather more language-balanced speech data.
- Includes over 10,000 hours for 22 languages and over 5,000 hours for 73 languages.
- Provides analyses of language coverage, audio quality, and transcription quality, plus baseline speech-recognition and neural-codec models.