AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
Ranking
Overall
88
Content
95
Popularity
71
Observed public metrics from 1 member.
Merged summary
TL;DR - AuK is an open-source foundation model that unifies instruction-driven speech generation and audio editing. Its distilled AuK-Flash variant enables four-step inference with a reported 4.5Ă— wall-clock speedup while retaining strong generation and editing performance.
- Training uses roughly 3.03 billion instruction–audio instances and 1.95 million hours of effective supervision across five generation, editing, enhancement, and separation task families.
- The architecture combines a multimodal LLM, a speech/audio/music VAE, and a hybrid rectified-flow Transformer for semantic and acoustic conditioning.
- Post-training applies human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation.
- The authors report leading zero-shot and instruction-controlled generation and editing results, competitive restoration performance, and release both code and model weights.
Sources (1)
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
Public signals
Hugging Face upvotes 164
TL;DR - AuK is an open-source foundation model that unifies instruction-driven speech generation and audio editing. Its distilled AuK-Flash variant enables four-step inference with a reported 4.5Ă— wall-clock speedup while retaining strong generation and editing performance.
- Training uses roughly 3.03 billion instruction–audio instances and 1.95 million hours of effective supervision across five generation, editing, enhancement, and separation task families.
- The architecture combines a multimodal LLM, a speech/audio/music VAE, and a hybrid rectified-flow Transformer for semantic and acoustic conditioning.
- Post-training applies human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation.
- The authors report leading zero-shot and instruction-controlled generation and editing results, competitive restoration performance, and release both code and model weights.