VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor
IBM
cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
Code: https://github.com/IBM/VAKRA
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 75/100
The gist: VAKRA (eValuating API and Knowledge Retrieval Agents) is a benchmark introduced in this paper for evaluating multi-hop reasoning across APIs and retrieval under tool-use policies in executable
Terminology
Summary
VAKRA (eValuating API and Knowledge Retrieval Agents) is a benchmark introduced in this paper for evaluating multi-hop reasoning across APIs and retrieval under tool-use policies in executable environments. It is built on the structured API generation pipeline of Elder et al. (2026) and transforms BIRD-SQL into more than 8,000 executable APIs derived from real databases across 62 domains. The benchmark extends this foundation with domain-aligned document collections, 2–5 hop reasoning chains that require agents to combine API interaction with document retrieval within a single trajectory, and a trajectory-level evaluation framework that re-executes predicted tool calls against live APIs, accommodating multiple valid paths.
Tasks are organized into three settings of increasing difficulty: (a) diverse API styles, (b) multi-hop reasoning over structured APIs, and (c) multi-source reasoning with natural-language tool-use policy constraints. The API styles setting includes three interaction styles: SLOT (compositional interfaces with 9 general-purpose tools), SEL (expanded function interfaces with 26 tools), and Dashboard APIs (endpoint-style interfaces with 116 tools per sample). The multi-hop reasoning setting requires 2–5 step reasoning chains over endpoint-style APIs, requiring entity disambiguation, parameter extraction from intermediate responses, and API schema alignment. The multi-source reasoning setting extends this to tasks requiring reasoning across structured APIs and unstructured document collections, in both single- and multi-turn settings, with a subset augmented with natural-language tool-use policies governing source selection.
The dataset statistics show the tuning split contains 710 samples for BI APIs (SEL), 614 for BI APIs (SLOT), 1,860 for Dashboard APIs, 346 for Multi-hop Reasoning, and 898 for Multi-source Multi-hop Reasoning. The test split contains 549, 1,397, 1,597, 869, and 644 samples respectively. Of the 664 samples in the multihop multisource dataset, 244 samples have policies. The multi-source hop-type distribution for the remaining 420 queries is included in Appendix C.2.
Correctness is verified by re-executing predicted tool calls against live APIs, accommodating multiple valid paths. The evaluation uses a waterfall mechanism with three stages: Stage 1 (Tool-Sequence Verification) re-executes each predicted tool call against the live environment and compares the set of tool responses against ground truth, using a programmatic containment check followed by an LLM-based judge adapted from CRAG for inconclusive cases; Stage 2 (Final Response Evaluation) uses an LLM judge to evaluate the agent's final response for groundedness in tool responses and answer correctness using the framework from Es et al. (2024); Stage 3 (Policy Adherence) verifies deterministically that no disallowed sources were consulted. GPT-OSS-120B is employed as the judge model for both Stage 1 and Stage 2 with temperature 0.
Using a fixed ReAct harness to isolate model capabilities from agent architecture, the paper evaluates frontier and open-weight models. The strongest model (GPT-5.5) achieves 70.4% on single-hop endpoint-style tasks but only 50–51% on compositional business-intelligence APIs. Most models lose over 50% accuracy as reasoning depth increases, and policy-constrained questions expose severe failures, with accuracy on unanswerable queries falling to 2.4% (Claude Opus 4.7). Trace analysis reveals that failures concentrate at language-mediated steps—entity disambiguation, cross-source grounding, and schema alignment—rather than tool invocation itself, indicating that compositional reasoning across heterogeneous sources remains the key bottleneck.
The paper's contributions are: (i) a tool-grounded benchmark with over 8,000 executable APIs across 62 domains, paired with domain-aligned document collections and trajectory-level evaluation that re-executes tool calls against live APIs; (ii) compositional multi-hop tasks with 2–5 hop chains combining structured API calls, cross-source retrieval, and natural-language tool-use policies; (iii) empirical analysis showing model rankings invert across API interaction paradigms, performance degrades sharply with reasoning depth, and tool-use policy adherence remains a critical weakness with failures localizing to language-mediated reasoning rather than tool invocation mechanics.
The human study for data quality assessment sampled 60 questions from each setting (869 Multi-hop Reasoning, 644 Multi-hop Multi-Source Reasoning) via stratified sampling across 12 semantic clusters, with three annotators independently scoring each question across 5 rubric dimensions—faithfulness, logical consistency, answer leakage, context sufficiency, and cross-source entity consistency—using a 1–4 ordinal scale. Inter-annotator agreement was 77% for Multi-hop Reasoning and 90% for Multi-hop Multi-Source Reasoning. Under the quality threshold (average score ≥ 3.0), 87% of Multi-hop Reasoning samples and 96% of Multi-hop Multi-Source Reasoning samples were rated as high-quality.
The paper concludes with three key findings. First, API interaction paradigm strongly influences difficulty: models that excel on endpoint-style APIs often struggle on compositional business-intelligence APIs, and vice versa, indicating that no single capability underlies tool-use proficiency. Second, multi-hop reasoning remains fragile—most models lose over 50% accuracy as reasoning depth increases, with failures concentrating at language-mediated steps (entity disambiguation, cross-source grounding, schema alignment) rather than tool invocation mechanics. Third, tool-use policy adherence is a critical weakness: when policies render questions unanswerable, even frontier models fail to recognize this, with accuracy falling as low as 2.4%. These results suggest that improving agentic systems requires advances in compositional reasoning and constraint interpretation, not merely better tool-calling interfaces.
Improvements for AI systems
Improvements to AI Systems:
- Add a
Policy-Aware Abstention
Module
-
Train models to detect when a natural-language tool-use policy makes a query unanswerable (e.g., disallowed data sources) and explicitly output
cannot answer
instead of guessing. -
Implement a dedicated fine-tuning objective on policy-violation examples, using the paper’s 244 policy-augmented samples to teach recognition of source restrictions before tool invocation.
-
Resulting capability: The system will refuse to answer or will request policy clarification when constraints conflict with the query, reducing hallucinated responses from 2.4% accuracy to near-zero on unanswerable queries.
- Enhance Cross-Source Grounding via Intermediate Entity Disambiguation
-
Add a pre-tool-call step that explicitly resolves entity mentions across structured APIs and unstructured documents (e.g., matching
Apple Inc.
in a document toAAPL
in a database) before generating tool calls. -
Use a contrastive learning objective on the paper’s multi-hop reasoning traces to align entity representations across modalities, with a separate verification layer that re-checks grounding after each retrieved document.
-
Resulting capability: The system will maintain consistent entity identity across heterogeneous sources, reducing failures at language-mediated steps (currently the primary bottleneck) and improving multi-hop accuracy by an estimated 20–30% relative.
- Implement Adaptive Reasoning-Depth Routing
-
Build a meta-controller that predicts the required reasoning depth (2–5 hops) from the query’s semantic complexity and dynamically switches between a fast single-hop path and a slower, more deliberate multi-hop path with intermediate verification.
-
Train the router on the paper’s hop-type distribution (from Appendix C.2) to recognize queries needing entity disambiguation, parameter extraction, or schema alignment, and allocate additional compute only when needed.
-
Resulting capability: The system will maintain high accuracy on deep reasoning chains (where models currently lose >50% accuracy) while preserving efficiency on simple tasks, achieving a better accuracy–latency tradeoff.
- Add a Schema-Alignment Preprocessor for Compositional APIs
-
Before executing tool calls on compositional business-intelligence APIs (SLOT/SEL styles), automatically map the query’s natural-language parameters to the correct API schema fields (e.g., converting
last quarter
to the correct date range parameter). -
Use the paper’s 8,000+ API definitions to train a schema-matching model that predicts field mappings from query text, with a fallback to retrieve the API documentation if confidence is low.
-
Resulting capability: The system will correctly invoke compositional APIs (where current models score only 50–51%) by reducing parameter-extraction errors, potentially closing the gap to endpoint-style performance.
- Implement a Trajectory-Level Self-Correction Loop
-
After each tool call, compare the live API response against a predicted response (using the paper’s Stage 1 verification logic) and, if mismatched, automatically re-plan the next tool call or retrieve additional documents before proceeding.
-
Integrate this as a reinforcement-learning reward signal during training, using the paper’s waterfall evaluation (tool-sequence verification + final response groundedness) as the reward function.
-
Resulting capability: The system will recover from intermediate errors (e.g., wrong parameter extraction) without human intervention, improving end-to-end task success on multi-hop and multi-source benchmarks by an estimated 15–25%.
- Add a
Paradigm-Aware
Model Ensemble
-
Train two specialized sub-models: one optimized for endpoint-style APIs (Dashboard) and one for compositional APIs (SLOT/SEL), then use a query classifier to route each task to the appropriate sub-model.
-
Use the paper’s finding that model rankings invert across paradigms to build a routing layer that selects the best-performing model per API style, rather than relying on a single generalist.
-
Resulting capability: The system will achieve near-best performance on both API styles simultaneously, avoiding the current tradeoff where a model excels on one paradigm but fails on the other.
Abstract
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e Valuating API and Knowledge Retrieval Agents), a benchmark of over 8, 000 executable APIs across 62 domains with tasks spanning three settings of increasing difficulty: diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source reasoning with natural-language tool-use policy constraints. Correctness is verified by re-executing predicted tool calls against live APIs, accommodating multiple valid paths. Using a fixed ReAct harness to isolate model capabilities from agent architecture, we evaluate frontier and open-weight models and find that even the best model achieves only 70.4% on single-hop endpoint-style tasks and drops to 50--51% on compositional APIs; performance degrades by over 50% as reasoning depth increases, and policy-constrained questions expose severe failures (as low as 2.4% on unanswerable queries). Trace analysis shows failures concentrate at language-mediated reasoning - entity disambiguation, cross-source grounding, rather than tool invocation mechanics. Code is available https://github.com/IBM/VAKRA. Dataset is available https://huggingface.co/datasets/ibm-research/VAKRA
Sources
- Automated Creation and Enrichment Framework for Improved Invocation of Enterprise APIs as Tools
- NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
- Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints
- ToolACE: Winning the Points of LLM Function Calling
- NesTools: A Dataset for Evaluating Nested Tool Learning Abilities of Large Language Models
- Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
- A New Query Expansion Approach via Agent-Mediated Dialogic Inquiry
- AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs
- RestGPT: Connecting Large Language Models with Real-World RESTful APIs
- gpt-oss-120b & gpt-oss-20b Model Card
- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- On the Tool Manipulation Capability of Open-source Large Language Models
- CRAG -- Comprehensive RAG Benchmark
- Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls
- WebArena: A Realistic Web Environment for Building Autonomous Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection