CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
summary
The gist
This research introduces chi-Bench, a novel benchmark designed to rigorously test the capabilities of AI agents in automating complex, end-to-end, long-horizon healthcare operations.
In short
$ ext{CHI}$-Bench tests AI agents' ability to automate complex, long-horizon healthcare tasks like prior authorization and utilization management. The benchmark stresses policy density and multi-role composition using a massive knowledge base. Findings show current agents struggle significantly with clinical reasoning and policy compliance, indicating a need for more robust grounding in complex medical rules.
Key concepts
- Policy Density
- This refers to the complexity of healthcare rules and policies that an AI agent must understand and apply correctly. It means the agent needs to navigate a vast library of detailed regulations, such as those found in managed-care handbooks, to make accurate decisions.
- Multi-role Composition
- This concept describes tasks where one goal requires the AI agent to switch between different functional personas or roles. For example, an agent might need to act as a patient advocate and a billing specialist simultaneously throughout a single process.
- In-Situ Verification
- This is a method used to judge if an AI's reasoning is correct by having another LLM check the agent's output against the source data. It ensures that the clinical decisions made by the agent are actually supported by cited policy sections and chart evidence.
- Long-horizon Tasks
- These are complex workflows in healthcare that require many sequential steps and long durations to complete, such as a full prior authorization process. Agents must maintain context and plan across numerous interactions to reach the final outcome.
Terminology used across episodes
This episode discusses
- CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? · Paper Radio
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- tau squared-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
- HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks
- Evaluating Large Language Models Trained on Code
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- GLM-5: from Vibe Coding to Agentic Engineering
- What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
- PubMedQA: A Dataset for Biomedical Research Question Answering
- MedCalc-Bench: Evaluating Large Language Models for Medical Calculations
- Kimi K2: Open Agentic Intelligence
- EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records
- The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Medchain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering
- Qwen3 Technical Report
- AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
- BioCoder: A Benchmark for Bioinformatics Code Generation with Large Language Models
The paper
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? · Read on arXiv
Haolin Chen, Deon Metelski, Leon Qi, Tao Xia, Joonyul Lee, Steve Brown, Kevin Riley, Frank Wang
actAVA.ai
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?".
Tom: Detailed Research Summary of chi-Bench: End-to-End Automation in Policy-Dense Healthcare Workflows This research introduces chi-Bench,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, we're starting with the title and the authors of this paper, "CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?"
Jane: The authors include a team from various institutions like Johns Hopkins Medicine and Stanford University. They’ve put together this benchmark to see if agents can manage the long, detailed tasks in healthcare workflows.
Lu: The title itself sets the stage perfectly because it focuses on end-to-end automation, which means the AI isn't just doing one small step; it has to complete a whole process from start to finish.
Meng: I wonder if they have enough real-world data to make these tests meaningful, or if they’re mostly relying on simulations.
Lalam: The focus on policy-rich workflows suggests this benchmark is pushing AI beyond simple pattern matching into true decision-making under complex constraints.
The paper's summary: Tom: Now let's look at what the paper actually summarizes about CHI-Bench, and I think the core idea is that they created a simulator with twenty healthcare applications to test agents across three specific domains.
Jane: They set up scenarios like Provider Prior Authorization, Payer Utilization Management, and RN Care Management to really stress-test these agents in different ways.
Lu: The key mechanism they introduce is grounding the agent's actions in a massive one thousand two hundred seventy-nine-document managed-care operations handbook that covers rules and policies.
Meng: So the AI has to use that huge rulebook to decide its next move across those twenty simulated apps while achieving a final status. That sounds like a lot of context switching at once.
Lalam: It really shows how crucial it is for agents to have access to structured knowledge, not just general world knowledge, when they are operating in an enterprise setting.
The paper's improvements: Tom: The authors point out a few key areas where they improved the approach, like introducing a specific skill architecture that connects the agent to that operations handbook more deeply.
Jane: They also added a verification layer, combining deterministic checks with an LLM judge to grade the final output and ensure it makes sense against the handbook.
Lu: I think one of the most significant improvements is enabling multi-role composition, which means the agent has to seamlessly switch between different specialized roles within a single task.
Meng: That ability to switch between, say, a provider role and then an insurance reviewer role in one go shows they're aiming for true operational realism rather than just sequential tool use.
Lalam: And they also look at how agents can interact in multi-turn dialogues, which is essential for things like patient outreach where you have to adapt your tone based on the conversation history.
Conclusion: Tom: So, wrapping up this discussion on "CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?", the paper shows that these agents are struggling when they have to handle the sheer complexity of policy and long sequences.
Jane: The main implication is that for AI to be truly useful in fields like healthcare administration, it needs to master grounding its decisions in detailed operational rules and manage multiple roles simultaneously.
Lu: It highlights a major gap between current agent capabilities and the high fidelity required for real enterprise workflows, especially in areas like Care Management where the reasoning is quite nuanced.
Meng: From an engineering standpoint, it suggests we need better ways to structure knowledge so that agents can navigate those policy landscapes without getting lost or making careless errors.
Lalam: I believe this research paves the way for a future where AI systems aren't just helpful tools, but reliable partners capable of handling deeply complex, multi-faceted operational tasks autonomously.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck