🛰️ Daily AI Frontier
‹ back to 2026-08-16

全球第一,碾压谷歌!中国版Thinking Machines诞生,语音赛道变天了

Industry & News Multimodal & Generative

Ranking

Overall 59
Content 70
Popularity 34

Observed public metrics from 1 member.

Representative image for 全球第一,碾压谷歌!中国版Thinking Machines诞生,语音赛道变天了

Merged summary

TL;DR - Chinese startup VUI Labs introduced Luna-TTS, a diffusion-based speech synthesis model claiming leading quality and latency results across public and internal benchmarks. Its block-diffusion architecture targets expressive, real-time voice agents.

  • Luna-TTS replaces token-by-token speech generation with bidirectional masked diffusion built from Qwen3-0.6B.
  • Luna-Codec separates semantic content from acoustic details across eight codebooks, aided by WavLM distillation.
  • Luna-TTS Realtime generates 1.28-second blocks and reports 41.6 ms first-packet latency and a 0.024 real-time factor on two H20 GPUs.
  • The system combines GRPO post-training, multilingual data, and explicit emotion and nonverbal tags for controllable speech.

Sources (1)

全球第一,碾压谷歌!中国版Thinking Machines诞生,语音赛道变天了

WeChat: 新智元 2026-08-14 arXiv:2608.11593
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-09 08:14:48.513656 UTC

TL;DR - Chinese startup VUI Labs introduced Luna-TTS, a diffusion-based speech synthesis model claiming leading quality and latency results across public and internal benchmarks. Its block-diffusion architecture targets expressive, real-time voice agents.

  • Luna-TTS replaces token-by-token speech generation with bidirectional masked diffusion built from Qwen3-0.6B.
  • Luna-Codec separates semantic content from acoustic details across eight codebooks, aided by WavLM distillation.
  • Luna-TTS Realtime generates 1.28-second blocks and reports 41.6 ms first-packet latency and a 0.024 real-time factor on two H20 GPUs.
  • The system combines GRPO post-training, multilingual data, and explicit emotion and nonverbal tags for controllable speech.
item →