MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents
TL;DR - MTVA-Bench evaluates the language model within cascaded voice agents under realistic multi-turn calling conditions, separating its decisions from ASR and TTS performance. Results show that overall quality depends far more on tool arguments, action ordering, rule compliance, and conversational behavior than on tool selection alone.
- Covers 49 agents, 490 reviewed scenarios, and seven languages, including transcript fragmentation and language/script requirements.
- Uses an LLM-simulated caller and a mock backend that responds to the tool arguments actually submitted.
- Combines deterministic tool-call checks with citation-based LLM judges for scenario rules and conversation quality, weighting task and conversation scores equally.
- Across seven models, correct-tool selection differed by only 6.4 points, while overall scores spanned 24.4 points.