🛰️ Daily AI Frontier
‹ back to 2026-09-20

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

arXiv cs.AI LLM Agents Pritish Mishra, Ishaan Kumar, Akshat Mandoli, Sudarshan Kamath 2026-09-17

TL;DR - MTVA-Bench evaluates the language model within cascaded voice agents under realistic multi-turn calling conditions, separating its decisions from ASR and TTS performance. Results show that overall quality depends far more on tool arguments, action ordering, rule compliance, and conversational behavior than on tool selection alone.

  • Covers 49 agents, 490 reviewed scenarios, and seven languages, including transcript fragmentation and language/script requirements.
  • Uses an LLM-simulated caller and a mock backend that responds to the tool arguments actually submitted.
  • Combines deterministic tool-call checks with citation-based LLM judges for scenario rules and conversation quality, weighting task and conversation scores equally.
  • Across seven models, correct-tool selection differed by only 6.4 points, while overall scores spanned 24.4 points.

view merged work →