全球第一,碾压谷歌!中国版Thinking Machines诞生,语音赛道变天了
TL;DR - Chinese startup VUI Labs introduced Luna-TTS, a diffusion-based speech synthesis model claiming leading quality and latency results across public and internal benchmarks. Its block-diffusion architecture targets expressive, real-time voice agents.
- Luna-TTS replaces token-by-token speech generation with bidirectional masked diffusion built from Qwen3-0.6B.
- Luna-Codec separates semantic content from acoustic details across eight codebooks, aided by WavLM distillation.
- Luna-TTS Realtime generates 1.28-second blocks and reports 41.6 ms first-packet latency and a 0.024 real-time factor on two H20 GPUs.
- The system combines GRPO post-training, multilingual data, and explicit emotion and nonverbal tags for controllable speech.