Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark
cs.AI, cs.DB
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/sparkcpark/synthetic_hospital
Terminology
Sources
- LongHealth: A Question Answering Benchmark with Long Clinical Documents
- Medical Large Language Model Benchmarks Should Prioritize Construct Validity
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
- HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks
- Generating Multi-label Discrete Patient Records using Generative Adversarial Networks
- TIMER: Temporal Instruction Modeling and Evaluation for Longitudinal Clinical Records
- AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator
- MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records
- Overview of the Problem List Summarization (ProbSum) 2023 Shared Task on Summarizing Patients' Active Diagnoses and Problems from Electronic Health Record Progress Notes
- Pattern Recognition or Medical Knowledge? The Problem with Multiple-Choice Questions in Medicine
- INSPECT: A Multimodal Dataset for Pulmonary Embolism Diagnosis and Prognosis
- MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
- People cannot distinguish GPT-4 from a human in a Turing test
- EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries
- Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes
- EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records
- FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
- Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents
- PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection