🛰️ Daily AI Frontier
‹ back to 2026-09-09

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

Research Multimodal & Generative

Ranking

Overall 88
Content 95
Popularity 71

Observed public metrics from 1 member.

Representative image for AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

Merged summary

TL;DR - AuK is an open-source foundation model that unifies instruction-driven speech generation and audio editing. Its distilled AuK-Flash variant enables four-step inference with a reported 4.5Ă— wall-clock speedup while retaining strong generation and editing performance.

  • Training uses roughly 3.03 billion instruction–audio instances and 1.95 million hours of effective supervision across five generation, editing, enhancement, and separation task families.
  • The architecture combines a multimodal LLM, a speech/audio/music VAE, and a hybrid rectified-flow Transformer for semantic and acoustic conditioning.
  • Post-training applies human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation.
  • The authors report leading zero-shot and instruction-controlled generation and editing results, competitive restoration performance, and release both code and model weights.

Sources (1)

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

arXiv cs.SD Ziyang Ma, Zhikang Niu, Wenming Tu, Tianrui Wang, Ruiqi Yan, Junxi Liu, Yanru Huo, Nickk Huang, Yang Liu, Qicong Xie, Zeyu Xie, Hui Wang, Haitao Li, Zixuan Jiang, Yalin Li, Jie Fang, Yifan Duan, Zeyue Tian, Guangzheng Li, Haina Zhu, Shuyi Wang, Jinwen Wang, Mingyu Cui, Tian Tan, Auden, Sen Liang, Steve Yves, Shan Yang, Liefeng Bo, Zilong Zheng, Kai Yu, Eng-Siong Chng, Xie Chen 2026-09-08 arXiv:2609.08936
Public signals Hugging Face upvotes 164
Providers: Hugging Face · Upvotes 164 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:22:30.637802 UTC

TL;DR - AuK is an open-source foundation model that unifies instruction-driven speech generation and audio editing. Its distilled AuK-Flash variant enables four-step inference with a reported 4.5Ă— wall-clock speedup while retaining strong generation and editing performance.

  • Training uses roughly 3.03 billion instruction–audio instances and 1.95 million hours of effective supervision across five generation, editing, enhancement, and separation task families.
  • The architecture combines a multimodal LLM, a speech/audio/music VAE, and a hybrid rectified-flow Transformer for semantic and acoustic conditioning.
  • Post-training applies human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation.
  • The authors report leading zero-shot and instruction-controlled generation and editing results, competitive restoration performance, and release both code and model weights.
item →