SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies
TL;DR - This paper presents a communication-efficient federated learning strategy for end-to-end SpeechLLM-based speech recognition. English and Italian case studies show competitive accuracy and stable decentralized training while reducing communication costs.
- Addresses high-dimensional parameters, gradient overhead, and distributed compute constraints.
- Compares federated and centralized training across varied acoustic conditions and speaking styles.
- Evaluates how speech encoder architectures affect federated English ASR performance.
- Establishes a foundation for privacy-preserving, multilingual SpeechLLM deployment.