🛰️ Daily AI Frontier
‹ back to 2026-08-13

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

Research LLM Agents

Ranking

Overall 80
Content 100
Popularity 34

Observed public metrics from 1 member.

Representative image for VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

Merged summary

TL;DR - VAKRA is a benchmark for evaluating agents that reason across structured APIs, document retrieval, and natural-language tool-use policies. Results show frontier models struggle sharply with compositional, multi-hop, and policy-constrained tasks.

  • Includes 8,000+ executable APIs across 62 domains and verifies predictions by replaying tool calls against live APIs.
  • The best model scores 70.4% on single-hop endpoint tasks but only 50–51% on compositional APIs.
  • Performance declines by over 50% as reasoning depth increases; accuracy reaches just 2.4% on some unanswerable, policy-constrained queries.
  • Failures center on entity disambiguation and cross-source grounding rather than tool-invocation mechanics.

Sources (1)

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

arXiv cs.AI Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor 2026-08-12 arXiv:2608.12282
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-09 08:17:05.675293 UTC

TL;DR - VAKRA is a benchmark for evaluating agents that reason across structured APIs, document retrieval, and natural-language tool-use policies. Results show frontier models struggle sharply with compositional, multi-hop, and policy-constrained tasks.

  • Includes 8,000+ executable APIs across 62 domains and verifies predictions by replaying tool calls against live APIs.
  • The best model scores 70.4% on single-hop endpoint tasks but only 50–51% on compositional APIs.
  • Performance declines by over 50% as reasoning depth increases; accuracy reaches just 2.4% on some unanswerable, policy-constrained queries.
  • Failures center on entity disambiguation and cross-source grounding rather than tool-invocation mechanics.
item →