长音频不丢词,行业词不用教,阿里发布Qwen-Audio-3.0-ASR-Flash
TL;DR - Alibaba released Qwen-Audio-3.0-ASR-Flash, an ASR model designed to improve long-audio consistency, specialized-term recognition, and custom hotword accuracy. It also converts speech directly into polished, structured text, reducing downstream processing.
- Long-audio context helps preserve names, terminology, and mixed Chinese-English expressions across segments.
- Specialized-term recall improved across internal evaluations, reaching 95.36% in medical scenarios.
- Hotword customization exceeded 90% accuracy on all test sets and 99% in most scenarios.
- The model removes filler words, resolves spoken corrections, and structures transcripts during recognition; it is available through Alibaba Cloud Model Studio.