Towards general auditory intelligence for machine listening and speaking
TL;DR - Wang et al. review progress toward general-purpose auditory intelligence spanning machine listening, speaking, speech interaction, and audio–visual understanding. The work highlights the convergence of these capabilities into more unified AI systems.
- Covers advances in both audio perception and speech generation.
- Examines speech-based human–machine interaction.
- Includes audio–visual understanding as a cross-modal capability.
- The provided abstract does not specify architectures, benchmarks, or quantitative results.