VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Ranking
Overall
80
Content
100
Popularity
34
Observed public metrics from 1 member.
Merged summary
TL;DR - VAKRA is a benchmark for evaluating agents that reason across structured APIs, document retrieval, and natural-language tool-use policies. Results show frontier models struggle sharply with compositional, multi-hop, and policy-constrained tasks.
- Includes 8,000+ executable APIs across 62 domains and verifies predictions by replaying tool calls against live APIs.
- The best model scores 70.4% on single-hop endpoint tasks but only 50–51% on compositional APIs.
- Performance declines by over 50% as reasoning depth increases; accuracy reaches just 2.4% on some unanswerable, policy-constrained queries.
- Failures center on entity disambiguation and cross-source grounding rather than tool-invocation mechanics.
Sources (1)
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - VAKRA is a benchmark for evaluating agents that reason across structured APIs, document retrieval, and natural-language tool-use policies. Results show frontier models struggle sharply with compositional, multi-hop, and policy-constrained tasks.
- Includes 8,000+ executable APIs across 62 domains and verifies predictions by replaying tool calls against live APIs.
- The best model scores 70.4% on single-hop endpoint tasks but only 50–51% on compositional APIs.
- Performance declines by over 50% as reasoning depth increases; accuracy reaches just 2.4% on some unanswerable, policy-constrained queries.
- Failures center on entity disambiguation and cross-source grounding rather than tool-invocation mechanics.