AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering
cs.CL, cs.AI
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/JohnnyNLP/agenthop
Terminology
Sources
- Cognitive Architectures for Language Agents
- ReAct: Synergizing Reasoning and Acting in Language Models
- AI Agents That Matter
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs
- Towards Scientific Intelligence: A Survey of LLM-based Scientific Agents
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- GAIA: a benchmark for General AI Assistants
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Kimi K2.5: Visual Agentic Intelligence
- Gemma 3 Technical Report
- Qwen3 Technical Report
- Reflexion: Language Agents with Verbal Reinforcement Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering