NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
TL;DR - NemotronLabs VoiceChat is an open, full-duplex speech-to-speech model that unifies real-time listening, transcription, reasoning, tool calling, and speech generation. It advances natural conversational interaction while showing that reliable tool arguments and end-to-end execution remain open challenges.
- Uses streaming speech and TTS components with parallel outputs for agent text and structured function calls.
- Achieves 100% takeover after user interruptions and resumes after backchannels in 93% of evaluated cases.
- Scores 55.1 normalized average on VoiceBench and 82.5% tool-selection F1 on Full-Duplex-Bench 3.0.
- Tool argument accuracy and end-to-end execution still need improvement.