🛰️ Daily AI Frontier
‹ back to 2026-08-13

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

arXiv cs.AI LLM Agents Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor 2026-08-12
Representative image for VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

TL;DR - VAKRA is a benchmark for evaluating agents that reason across structured APIs, document retrieval, and natural-language tool-use policies. Results show frontier models struggle sharply with compositional, multi-hop, and policy-constrained tasks.

  • Includes 8,000+ executable APIs across 62 domains and verifies predictions by replaying tool calls against live APIs.
  • The best model scores 70.4% on single-hop endpoint tasks but only 50–51% on compositional APIs.
  • Performance declines by over 50% as reasoning depth increases; accuracy reaches just 2.4% on some unanswerable, policy-constrained queries.
  • Failures center on entity disambiguation and cross-source grounding rather than tool-invocation mechanics.

view merged work →