VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
TL;DR - VAKRA is a benchmark for evaluating agents that reason across structured APIs, document retrieval, and natural-language tool-use policies. Results show frontier models struggle sharply with compositional, multi-hop, and policy-constrained tasks.
- Includes 8,000+ executable APIs across 62 domains and verifies predictions by replaying tool calls against live APIs.
- The best model scores 70.4% on single-hop endpoint tasks but only 50–51% on compositional APIs.
- Performance declines by over 50% as reasoning depth increases; accuracy reaches just 2.4% on some unanswerable, policy-constrained queries.
- Failures center on entity disambiguation and cross-source grounding rather than tool-invocation mechanics.