NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
Ranking
Overall
79
Content
95
Popularity
43
Observed public metrics from 1 member.
Merged summary
TL;DR - NemotronLabs VoiceChat is an open, full-duplex speech-to-speech model that unifies real-time listening, transcription, reasoning, tool calling, and speech generation. It advances natural conversational interaction while showing that reliable tool arguments and end-to-end execution remain open challenges.
- Uses streaming speech and TTS components with parallel outputs for agent text and structured function calls.
- Achieves 100% takeover after user interruptions and resumes after backchannels in 93% of evaluated cases.
- Scores 55.1 normalized average on VoiceBench and 82.5% tool-selection F1 on Full-Duplex-Bench 3.0.
- Tool argument accuracy and end-to-end execution still need improvement.
Sources (1)
NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
Public signals
Hugging Face upvotes 0
TL;DR - NemotronLabs VoiceChat is an open, full-duplex speech-to-speech model that unifies real-time listening, transcription, reasoning, tool calling, and speech generation. It advances natural conversational interaction while showing that reliable tool arguments and end-to-end execution remain open challenges.
- Uses streaming speech and TTS components with parallel outputs for agent text and structured function calls.
- Achieves 100% takeover after user interruptions and resumes after backchannels in 93% of evaluated cases.
- Scores 55.1 normalized average on VoiceBench and 82.5% tool-selection F1 on Full-Duplex-Bench 3.0.
- Tool argument accuracy and end-to-end execution still need improvement.