GPAgentBench-2K: Benchmarking Large Language Model Agents in Complex Clinical Action Space
cs.CL
Submitted: 2026-08-31
Updated: 2026-09-02
Code: https://github.com/jianing-lab/GPAgentBench
Terminology
Sources
- Real-World Doctor Agent with Proactive Consultation through Multi-Agent Reinforcement Learning
- Constrained Group Relative Policy Optimization
- Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
- ART: Action-based Reasoning Task Benchmarking for Medical AI Agents
- MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems
- AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
- MedGemma Technical Report
- Qwen3 Technical Report
- MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering