CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?".
Tom: Detailed Research Summary of chi-Bench: End-to-End Automation in Policy-Dense Healthcare Workflows This research introduces chi-Bench,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, we're starting with the title and the authors of this paper, "CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?"
Jane: The authors include a team from various institutions like Johns Hopkins Medicine and Stanford University. They’ve put together this benchmark to see if agents can manage the long, detailed tasks in healthcare workflows.
Lu: The title itself sets the stage perfectly because it focuses on end-to-end automation, which means the AI isn't just doing one small step; it has to complete a whole process from start to finish.
Meng: I wonder if they have enough real-world data to make these tests meaningful, or if they’re mostly relying on simulations.
Lalam: The focus on policy-rich workflows suggests this benchmark is pushing AI beyond simple pattern matching into true decision-making under complex constraints.
The paper's summary: Tom: Now let's look at what the paper actually summarizes about CHI-Bench, and I think the core idea is that they created a simulator with twenty healthcare applications to test agents across three specific domains.
Jane: They set up scenarios like Provider Prior Authorization, Payer Utilization Management, and RN Care Management to really stress-test these agents in different ways.
Lu: The key mechanism they introduce is grounding the agent's actions in a massive one thousand two hundred seventy-nine-document managed-care operations handbook that covers rules and policies.
Meng: So the AI has to use that huge rulebook to decide its next move across those twenty simulated apps while achieving a final status. That sounds like a lot of context switching at once.
Lalam: It really shows how crucial it is for agents to have access to structured knowledge, not just general world knowledge, when they are operating in an enterprise setting.
The paper's improvements: Tom: The authors point out a few key areas where they improved the approach, like introducing a specific skill architecture that connects the agent to that operations handbook more deeply.
Jane: They also added a verification layer, combining deterministic checks with an LLM judge to grade the final output and ensure it makes sense against the handbook.
Lu: I think one of the most significant improvements is enabling multi-role composition, which means the agent has to seamlessly switch between different specialized roles within a single task.
Meng: That ability to switch between, say, a provider role and then an insurance reviewer role in one go shows they're aiming for true operational realism rather than just sequential tool use.
Lalam: And they also look at how agents can interact in multi-turn dialogues, which is essential for things like patient outreach where you have to adapt your tone based on the conversation history.
Conclusion: Tom: So, wrapping up this discussion on "CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?", the paper shows that these agents are struggling when they have to handle the sheer complexity of policy and long sequences.
Jane: The main implication is that for AI to be truly useful in fields like healthcare administration, it needs to master grounding its decisions in detailed operational rules and manage multiple roles simultaneously.
Lu: It highlights a major gap between current agent capabilities and the high fidelity required for real enterprise workflows, especially in areas like Care Management where the reasoning is quite nuanced.
Meng: From an engineering standpoint, it suggests we need better ways to structure knowledge so that agents can navigate those policy landscapes without getting lost or making careless errors.
Lalam: I believe this research paves the way for a future where AI systems aren't just helpful tools, but reliable partners capable of handling deeply complex, multi-faceted operational tasks autonomously.
Haolin Chen, Deon Metelski, Leon Qi, Tao Xia, Joonyul Lee, Steve Brown, Kevin Riley, Frank Wang
actAVA.ai
cs.CL, cs.AI
Submitted: 2026-05-15
Updated: 2026-09-30
Comments: Website: https://actava.ai/benchmarks Code: https://github.com/actava-ai/chi-bench Dataset: https://huggingface.co/datasets/actava/chi-bench
Code: https://github.com/anthropics/claude-code
Project page: https://actava.ai/benchmarks
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 80/100
The gist: This research introduces chi-Bench, a novel benchmark designed to rigorously test the capabilities of AI agents in automating complex, end-to-end, long-horizon healthcare operations.
Key concepts
- Policy Density
- This refers to the complexity of healthcare rules and policies that an AI agent must understand and apply correctly. It means the agent needs to navigate a vast library of detailed regulations, such as those found in managed-care handbooks, to make accurate decisions.
- Multi-role Composition
- This concept describes tasks where one goal requires the AI agent to switch between different functional personas or roles. For example, an agent might need to act as a patient advocate and a billing specialist simultaneously throughout a single process.
- In-Situ Verification
- This is a method used to judge if an AI's reasoning is correct by having another LLM check the agent's output against the source data. It ensures that the clinical decisions made by the agent are actually supported by cited policy sections and chart evidence.
- Long-horizon Tasks
- These are complex workflows in healthcare that require many sequential steps and long durations to complete, such as a full prior authorization process. Agents must maintain context and plan across numerous interactions to reach the final outcome.
Terminology
Summary
This research introduces ** chi-Bench**, a novel benchmark designed to rigorously test the capabilities of AI agents in automating complex, end-to-end, long-horizon healthcare operations. The core premise is that current benchmarks fail to adequately stress three critical capabilities essential for real-world enterprise automation: policy density (requiring grounding in vast rule libraries), multi-role composition (where a single task necessitates playing multiple distinct roles with necessary handoffs), and multilateral interaction (involving multi-turn dialogues like peer reviews and patient outreach).
chi-Bench evaluates AI agents across three distinct, long-horizon healthcare domains:
-
Provider Prior Authorization (PA): Involves verifying coverage, gathering evidence, submitting the initial packet, and managing subsequent responses (RFIs, peer-to-peer reviews, appeals).
-
Payer Utilization Management (UM): Requires intake of requests, checking plan policies against them, escalating through nurse and physician reviewers for determination issuance.
-
RN Care Management (CM): Involves reviewing clinical charts, contacting the patient directly to administer assessments, and managing ongoing care plans.
The benchmark is constructed to satisfy nine specific evaluation axes: Healthcare Workflows (HC), API tools (API tools), Long-horizon tasks (Long-horiz.), Hidden state, In-Situ verification (In-Situ LLM judge), and a comprehensive LLM judge.
Data Grounding: The agent's decisions are strictly guided by a massive knowledge base—specifically, a 1,279-document Managed-Care Operations Handbook alongside an extensive medical and insurance policy library. Crucially, the simulator exposes high-fidelity interactions across 20 healthcare applications via 87 Multi-Context Protocol (MCP) tools, ensuring that the agent must drive these tools to achieve a terminal status through tool calls and artifact generation.
The evaluation is designed with a robust, two-layer verification system:
-
Deterministic Layer: Reads the simulator’s persisted world state and event log to verify terminal case status, correct routing, structured payloads, and the presence of required side-effect artifacts.
-
LLM Judge Layer: Reads the same records to grade whether clinical reasoning is accurately grounded in cited policy sections by name, whether criterion evaluations are internally consistent with chart evidence, and if the final determination logically follows from the recorded rationale.
Key Performance Metrics:
-
Pass@1 (pass@1): Measures the percentage of tasks successfully completed on a single attempt.
-
** pass cubed Reliability:** A stricter metric requiring success on every individual case across multiple trials, designed to capture operational robustness. The paper notes that executing all tasks in a single session slumps performance significantly (to 3.8%), highlighting the need for reliable, repeatable performance.
The experiment tested 30 different agent harness/model configurations. The results reveal significant limitations in current agent capabilities:
-
Overall Performance: The strongest configuration (Claude Code + Opus 4.6) resolves only **28.0% of tasks at pass@1 **, and the pass cubed reliability metric collapses, indicating severe run-to-run inconsistency that is unacceptable for production deployment.
-
Failure Categorization: Failures are systematically categorized into three first-level modes:
-
Clinical Reasoning (35.4%): The single largest bottleneck, suggesting agents struggle with the nuanced interpretation of clinical situations and policy application.
-
Workflow Completion (23.3%): Failures related to skipping required steps or failing to complete the necessary sequence of actions within the long-horizon task structure.
-
Policy Compliance (13.2%): Failures stemming from incorrect application or misreading of complex policy criteria, which is further detailed at the second level (e.g., criteria misapplication, skipped steps).
Skill Invocation Analysis: A key finding relates to knowledge access: the mean policy-read recall across all trials was only 29.3%. While this low recall suggests agents are not accessing the full scope of the handbook, a strong positive correlation (r = +0.77) exists between cell-mean policy-read recall and pass@1, confirming that accessing the handbook is a necessary signal, though it is insufficient on its own.
Domain-Specific Bottlenecks:
- Care Management (CM): This domain exhibits the highest failure rate in Clinical Reasoning (50.
Improvements for AI systems
Based on the scientific paper χ-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
, here are specific improvements for AI systems and what those improvements enable them to do:
) 1. Implement a Hierarchical Skill Architecture grounded in a Structured Knowledge Base (The Managed-Care Operations Handbook Skill
).
Improvement: Instead of relying solely on large language model (LLM) context or prompt injection, integrate domain knowledge via structured skills organized as a progressive-disclosure manual (as shown in Figure 7). This skill must explicitly route the agent to role-specific sub-skills (PA specialist, UM reviewer, care manager) and access shared medical/platform libraries.
What it enables:
The AI system can navigate complex, policy-dense healthcare workflows with high fidelity. It ensures decisions are grounded in explicit rules (policy retrieval) rather than relying on generalized medical knowledge or hallucination. This drastically reduces
Policy-Compliancefailures by ensuring every action is traceable to a cited handbook section, improving reliability from 28% (pass@1) to a higher level of adherence.
) 2. Adopt In-Situ Verification with a Multi-Layered Verifier (Deterministic Contract + Rubric LLM Judge).
Improvement: The agent's output must be validated not just by the LLM judge, but by comparing the resulting world state, event log, and artifacts against a deterministic contract. The LLM judge must then grade this persisted state against rubric-based criteria that require explicit policy citations.
What it enables:
The system gains high reliability in irreversible workflows. It prevents
Workflow-Completionfailures by ensuring terminal actions are correctly invoked and preventsPolicy-Complianceerrors by forcing the agent to reference specific policy sections, mitigating the risk of substituting general knowledge for actual policy adherence.
) 3. Enable Multi-Role Composition with Terminal Handoffs in Single Tasks.
Improvement: Design agents to seamlessly switch between distinct, role-scoped tool surfaces (e.g., Provider Clinician → Payer UM Nurse → Medical Director) within a single, long-horizon task without losing context or being able to edit previous steps.
What it enables:
The system can execute complex, multi-stage enterprise tasks end-to-end in one session. It allows an agent to simulate a complete clinical review process (PA/UM/CM) by correctly switching roles, handling handoffs as terminal events, and managing the accumulated state across all required domains.
) 4. Integrate Multilateral Interaction Capabilities for Live Dialog.
Improvement: Equip agents with the ability to shift from background tool execution to live, multi-turn dialogs with simulated parties (e.g., patient outreach in Care Management, peer-to-peer review in UM). The agent must be able to collect information incrementally from these human interactions and carry those results back into the formal workflow.
What it enables:
The system can handle dynamic, real-world engagement. In Care Management, it can adapt its communication style based on patient consent profiles (Engaged vs. Refusing) and successfully conduct a structured assessment interview by adhering to patient-centered techniques like
consent triggers,preventing premature or inappropriate clinical disclosures.
) 5. Develop Adaptive Patient Persona Modeling with Consent Triggers and Anti-Triggers.
Improvement: The agent must maintain an internal, opaque model of the patient's trust state, defined by explicit consent triggers (actions that build rapport) and anti-triggers (actions that damage rapport). Communication strategy must dynamically adapt based on these inferred profiles.
What it enables:
The system can execute sensitive clinical outreach ethically. It prevents
illegitimate consentfailures by ensuring the agent doesn't bypass required assessment steps prematurely, instead waiting for specific patient cues, thereby creating a functional prerequisite for administering structured assessments like PHQ-9 or GAD-7 scores.
) 6. Implement Policy-Grounded Task Generation via Constraint Sampling.
Improvement: When generating candidate tasks (Step 1 of the task construction pipeline), use a rejection-sampling loop around the LLM agent generator, constrained by serialized slices of relevant system state graphs (e.g., payer policy decision trees) and structured schemas that require every output field to be anchored to a citation in the handbook or policy.
What it enables:
The system can generate realistic, high-quality training data for agents that are inherently policy-aware. This ensures that the agent learns to only emit actions supported by the specific policy context provided, drastically reducing
Policy-Complianceerrors during deployment.
Abstract
End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions must be grounded in a large library of medical, insurance, and operational rules; Multi-role composition: a single task requires the agent to play multiple roles with handoffs; and multilateral interaction: intermediate workflow steps are multi-turn dialogs, such as peer-to-peer review and patient outreach. We introduce χ-Bench, a benchmark of long-horizon healthcare workflows across three domains: provider prior authorization, payer utilization management, and care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed via 87 MCP tools, which it must drive to a terminal status through tool calls and writing the role's artifacts, guided by a 1,290+ document managed-care operations handbook skill. Across 30 agent harness/models configurations, the best agent resolves only 28.0% of tasks, no agent clears 20% on strict pass cubed, and executing all tasks in a single session slumps the performance to 3.8%. These results raise the hypothesis that similar gaps are likely to surface in other policy-dense, role-composed, irreversible enterprise domains.
Sources
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
- HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks
- Evaluating Large Language Models Trained on Code
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- GLM-5: from Vibe Coding to Agentic Engineering
- What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
- PubMedQA: A Dataset for Biomedical Research Question Answering
- MedCalc-Bench: Evaluating Large Language Models for Medical Calculations
- Kimi K2: Open Agentic Intelligence
- EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records
- The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Medchain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering
- Qwen3 Technical Report
- AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
- BioCoder: A Benchmark for Bioinformatics Code Generation with Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering