🛰️ Daily AI Frontier
‹ back to 2026-09-21

NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

arXiv cs.CL Multimodal & Generative Jagadeesh Balam, Travis Bartley, Edresson Casanova, Sanjay Chauhan, Chen Chen, Zhehuai Chen, Zijia Chen, Francesco Ciannella, Slyne Deng, Mikyas Desta, Harishchandra Dubey, Slim Essid, Nourchene Ferchichi, Boris Ginsburg, Mariana Graterol Fuenmayor, Negar Habibi, Kevin Hu, Anand Joseph, Viraj Karandikar, Myungjong Kim, Viacheslav Klimkov, Seelan Lakshmi Narasimhan, Lily Lee, Jason Li, Eileen Long, Ameya Mahabaleshwarkar, Aditya Malte, Adi Margolin, Sasha Meister, Valentin Mendelev, Oluwatobi Olabiyi, Ankita Pasad, Yifan Peng, Elena Rastorgueva, Jayda Ritchie, Jason Roche, Nikhil Srihari, Yuanhang Su, Yoshi Suhara, Viet Anh Trinh, Jinhan Wang, Piotr Zelasko, Hui Wang, Puhui Meng, Chaosen Zhang, Yunsheng Liu, Shawn Wang, Wenjing Li, Zhonglei He 2026-09-18

TL;DR - NemotronLabs VoiceChat is an open, full-duplex speech-to-speech model that unifies real-time listening, transcription, reasoning, tool calling, and speech generation. It advances natural conversational interaction while showing that reliable tool arguments and end-to-end execution remain open challenges.

  • Uses streaming speech and TTS components with parallel outputs for agent text and structured function calls.
  • Achieves 100% takeover after user interruptions and resumes after backchannels in 93% of evaluated cases.
  • Scores 55.1 normalized average on VoiceBench and 82.5% tool-selection F1 on Full-Duplex-Bench 3.0.
  • Tool argument accuracy and end-to-end execution still need improvement.

view merged work →