ClinTraceBench: Source-Verifiable Longitudinal Clinical Reasoning over EHR-Derived Dialogues
cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: Findings of EMNLP 2026
Code: https://github.com/HathyHuimin/ClinTraceBench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Clinical LLM assistants must reason over multi-visit patient trajectories, yet whether the compact history representations used to scale them---retrieval, structured timelines, LLM summaries, agentic
Terminology
Abstract
Clinical LLM assistants must reason over multi-visit patient trajectories, yet whether the compact history representations used to scale them---retrieval, structured timelines, LLM summaries, agentic memory---preserve the longitudinal signal clinical reasoning needs has not been measured. We introduce ClinTraceBench: 385 MIMIC-IV-derived verified dialogues with event-ID provenance, a nine-task taxonomy (T1--T9), and L0--L4 deterministic + L5 human-audit validation (98.92% agreement). We evaluate eight history representation strategies---a no-context floor, last-visit-only, full-context, BGE-M3 dense-retrieval, two compression schemes, and two agentic-memory systems (Mem0, A-Mem)---across four backbones (DeepSeek-V3, GPT-4o-mini, Haiku 4.5, Sonnet 4.6) on 6, 271 questions: 32 cells, 200, 672 predictions. Four findings: (SP4) a controlled T3 injection probe isolates compression-induced relation loss---with the attribution sentence present before construction, Mem0, A-Mem and llm-summary still recover only 0--5.3% of the injected positives; (SP1) compressed strategies pay an aggregation tax on multi-visit trends and cross-patient comparisons; (SP2) the blind-to-full gap spans +29.8 pp (GPT-4o-mini) to +62.7 pp (Haiku); (SP3) abstention scales non-monotonically with context length. On the Pareto frontier Haiku dominates Sonnet under full-context (25.76 vs. 106.21), inverting the ``biggest backbone wins'' heuristic.
Sources
- Language Models (Mostly) Know What They Know
- MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
- Large Language Models with Temporal Reasoning for Longitudinal Clinical Summarization and Prediction
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- MemGPT: Towards LLMs as Operating Systems
- A-MEM: Agentic Memory for LLM Agents
- A Survey on the Memory Mechanism of Large Language Model based Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering