Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores
cs.CL, cs.AI
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/nicoveraz/calcbounds
Terminology
Sources
- MedCalc-Bench Doesn't Measure What You Think: A Benchmark Audit and the Case for Open-Book Evaluation
- AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
- Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering