Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models
cs.LG, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-28
Comments: Submitted to ML4H 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: How large language models (LLMs) integrate patient risk with clinical cost tradeoffs remains poorly understood.
Terminology
Abstract
How large language models (LLMs) integrate patient risk with clinical cost tradeoffs remains poorly understood. We investigated how four open-weight LLMs (Qwen-2.5-7B/32B and Llama-3.1-8B/70B) internally represent cost tradeoffs, how these representations relate to clinical predictions, and whether decisions shift as predicted by the specified cost direction and magnitude. Using a public diabetes dataset, we varied 11 false-negative (FN) to false-positive (FP) cost ratios across three phrasings and examined representations and behavioral outputs. Patient risk was linearly recoverable on par with conventional classifiers (AUC about 0.83), and cost direction was recoverable in every model. However, representational shifts in cost direction tracked output changes only in the two larger models, and responses to cost magnitude were predominantly direction-agnostic. Only 2 of 12 model-phrasings showed both opposing responses to increasing FN versus FP costs and cost-correct ordering. Representationally, a direction fitted on one cost side did not invert when transferred to the other, as expected under mirror-symmetric encoding. These findings suggest that LLMs encode risk and cost information but do not reliably integrate them into cost-correct decisions. Clinical evaluations should therefore include tradeoff tests, phrasing sensitivity, and default operating points alongside predictive performance.
Sources
- Understanding intermediate layers using linear classifier probes
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Mechanistic Interpretability for AI Safety -- A Review
- Finding Neurons in a Haystack: Case Studies with Sparse Probing
- No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- Attention Is All You Need
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks