🛰️ Daily AI Frontier
‹ back to 2026-09-20

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

Research LLM Agents

Ranking

Overall 78
Content 90
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Merged summary

TL;DR - MTVA-Bench evaluates the language model within cascaded voice agents under realistic multi-turn calling conditions, separating its decisions from ASR and TTS performance. Results show that overall quality depends far more on tool arguments, action ordering, rule compliance, and conversational behavior than on tool selection alone.

  • Covers 49 agents, 490 reviewed scenarios, and seven languages, including transcript fragmentation and language/script requirements.
  • Uses an LLM-simulated caller and a mock backend that responds to the tool arguments actually submitted.
  • Combines deterministic tool-call checks with citation-based LLM judges for scenario rules and conversation quality, weighting task and conversation scores equally.
  • Across seven models, correct-tool selection differed by only 6.4 points, while overall scores spanned 24.4 points.

Sources (1)

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

arXiv cs.AI Pritish Mishra, Ishaan Kumar, Akshat Mandoli, Sudarshan Kamath 2026-09-17 arXiv:2609.20152
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:15:09.078667 UTC

TL;DR - MTVA-Bench evaluates the language model within cascaded voice agents under realistic multi-turn calling conditions, separating its decisions from ASR and TTS performance. Results show that overall quality depends far more on tool arguments, action ordering, rule compliance, and conversational behavior than on tool selection alone.

  • Covers 49 agents, 490 reviewed scenarios, and seven languages, including transcript fragmentation and language/script requirements.
  • Uses an LLM-simulated caller and a mock backend that responds to the tool arguments actually submitted.
  • Combines deterministic tool-call checks with citation-based LLM judges for scenario rules and conversation quality, weighting task and conversation scores equally.
  • Across seven models, correct-tool selection differed by only 6.4 points, while overall scores spanned 24.4 points.
item →