Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance
cs.CL, cs.AI
Submitted: 2026-06-17
Updated: 2026-08-31
Code: https://github.com/qgpmztmf/PhysAssistBench
Terminology
Sources
- Synthetic Patient-Physician Dialogue Generation from Clinical Notes Using LLM
- GLM-5: from Vibe Coding to Agentic Engineering
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Kimi K2: Open Agentic Intelligence
- Seed1.8 Model Card: Towards Generalized Real-World Agency
- ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?
- EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries
- LLMs Get Lost In Multi-Turn Conversation
- FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
- ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- Qwen3 Technical Report
- Qwen3.5-Omni Technical Report
- RCBSF: A Multi-Agent Framework for Automated Contract Revision via Stackelberg Game
- OpenAI GPT-5 System Card
- Benchmarking LLM Tool-Use in the Wild
- Towards Expert-Level Medical Question Answering with Large Language Models
- MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering