Single-turn emergency psychiatric triage across 15 frontier AI chatbots

arXiv:2604.25415 · q-bio.NC, cs.AI, cs.HC · Submitted 2026-04-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: I'm Ines, and with me are Marcus and Yuki, guest researcher.

Marcus: Today's paper: "Single-turn emergency psychiatric triage across 15 frontier AI chatbots".

Ines: Frontier AI chatbots are increasingly used for health advice, but their performance in psychiatric triage remains undercharacterized, making it crucial to understand how these models handle urgent mental health disclosures.

Marcus: First, who's behind it and why it matters.

Paper summary: Ines: So, wrapping up, the paper "Single-turn emergency psychiatric triage across fifteen frontier AI chatbots" points out that while these models are quite accurate for truly critical emergencies when they have all the facts, they consistently show a tendency to over-triage situations that would only require routine or intermediate care.

Marcus: The authors are essentially highlighting a calibration gap in how these LLMs assess urgency in psychiatric contexts, showing that the system isn't perfectly tuned across the entire spectrum of risk levels. They found emergency under-triage was rare at five point six percent, but the overall error pattern leaned toward over-caution.

Yuki: From a broader perspective, this research suggests that as we integrate AI into medical decision support systems, especially for sensitive areas like mental health, we can't just rely on high overall accuracy figures without looking closely at where those errors are concentrated in terms of severity.

Ines: It speaks to the fact that the structure of a single-turn query doesn't capture all the complexity of real patient interactions, which is a limitation they pointed out regarding synthetic vignettes versus actual noisy disclosures.

Marcus: And their conclusion about the implication is that this pattern might stem from commercial pressures or risk aversion built into how these models are trained, leading them to err on the side of caution in response to ambiguous inputs.

Yuki: If this tendency toward over-caution is encoded in the model weights, it means that as AI becomes more prevalent in public health advice, we need robust methods to ensure those safety-oriented post-training procedures don't just bake a systemic bias into how they handle ambiguity.

Conclusion: Ines: So, to wrap up this discussion, we've seen how these frontier AI chatbots struggle when it comes to accurately prioritizing psychiatric emergencies based on a single user input.

Marcus: I think the authors are focusing on a really specific problem there, trying to pinpoint exactly where the models get confused when making triage decisions.

Yuki: It's interesting how they structured their benchmark using those one hundred twelve vignettes, which gives us a clear picture of what kind of clinical scenarios they are testing against.

Ines: Exactly, and looking at the title, "Single-turn emergency psychiatric triage across fifteen frontier AI chatbots," it tells us the scope is quite broad with a lot of different models involved.

Marcus: That breadth is key; seeing results across fifteen different systems really shows us if this behavior is a consistent pattern or just an outlier in one specific architecture.

Yuki: And thinking about the authors, they're clearly looking at the practical application here, trying to understand the real-world impact of these tools on patient care pathways.

Ines: Their main conclusion seems to be that while some models are decent at catching true emergencies when everything is provided, they have a significant tendency toward over-triage for lower-risk situations.

Marcus: That suggests there's a calibration issue where the AI leans toward caution even when the presented data points to less urgent needs.

Yuki: From my perspective, this speaks to how we’re training these models; if they are being pushed too hard on risk aversion during post-training, it could affect their real-world utility in nuanced situations.

Ines: It really raises questions about what's actually happening inside those model weights that leads to this systematic over-cautiousness.

Marcus: And the implication for us is that we need to look closely at how these safety procedures are influencing the models' decision-making process under pressure.

Yuki: So, as we move forward, I think it’s crucial for us to investigate if this tendency to over-triage generalizes when users give them more complex, fragmented information in a continuous conversation.

Veith Weilnhammer, *, Lennart Luettgau, *, Christopher Summerfield, %*, Viknesh Sounderajah&,, & Elise Wilkinson', ' Virginia Corno', Matthew M Nour&,, '

Max Planck UCL Centre for Computational Psychiatry and Ageing Research, London, UK · $ UK AI Security Institute · % Department of Experimental Psychology, University of Oxford, Oxford, UK · & Microsoft AI, London, UK

q-bio.NC, cs.AI, cs.HC

Submitted: 2026-04-28

Updated: 2026-09-30

Code: https://github.com/veithweilnhammer/chatbot-one-shot-psychiatric-triage

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: Frontier AI chatbots are increasingly used for health advice, but their performance in psychiatric triage remains undercharacterized, making it crucial to understand how these models handle urgent

Key concepts

Psychiatric Vignettes
These are simulated patient scenarios created by clinicians that combine specific psychiatric conditions (like psychosis) with defined risk factors (like suicidal ideation). They serve as the input for testing how well an AI can correctly categorize a mental health disclosure into one of four urgency levels.
Triage Levels (A-D)
These are four predefined urgency categories used to label psychiatric situations. Level D represents an immediate emergency requiring care now, while Level A is routine. The study measures the model's ability to correctly assign these labels based on the provided information.
Over-triage Bias
This refers to a systematic tendency for the AI to classify a presentation as more urgent than it actually is. The research found that models were significantly more likely to over-triage lower and intermediate risk cases, indicating an inherent caution in their decision-making process.

Terminology

Summary

Frontier AI chatbots are increasingly used for health advice, but their performance in psychiatric triage remains undercharacterized, making it crucial to understand how these models handle urgent mental health disclosures. The core finding of this study is that while frontier AI chatbots can recognize high-acuity psychiatric emergencies with near-zero error rates when presented with complete information, they exhibit a marked pattern of over-triage for clinical presentations carrying low and intermediate risk levels.

How it works

The study utilized a benchmark of 112 psychiatric vignettes, each paired with one of four original triage labels (A: routine; B: assessment within 1 week; C: assessment within 24 to 48 hours; D: emergency care now). These vignettes were constructed by combining nine psychiatric presentation clusters (such as suicidality or primary psychosis) with nine focal risk dimensions (such as suicidal risk or self-neglect), resulting in 28 clinically authored presentation-by-risk groups. Each group contributed four distinct clinical vignettes, one at each triage level.

Data Generation and Ground Truth

The creation of the dataset involved a multi-step pipeline. First, a clinician-validated LLM (Claude Opus 4.6) was instructed to construct a vignette with sufficient diagnostic and contextual detail to support the benchmark triage judgment from text alone, alongside explicit clinical reasoning for why it met or failed adjacent urgency thresholds. Second, an independent instance of Opus 4.6 was used to rewrite this summary as a simulated single-message disclosure suitable for a user query in a conversation. These simulated messages had a median word count of 564 and were either self-reports (76%) or collateral reports from another person (24%).

Evaluation Metrics and Outcomes

The primary outcome measure was the under-triage rate for emergency (level D) trials, which showed that emergency under-triage occurred in only 23 of 410 emergency (level D) trials (5.6%) across all target models. Secondary outcomes included level-specific and overall accuracy, mean signed ordinal error, and mean absolute ordinal error. The study found that accuracy was highest for level D vignettes (94.3%) and lowest for level B vignettes (19.7%). Furthermore, the mean signed ordinal error was positive (+0.47 triage levels), indicating net over-triage, with dispersion being highest around the middle triage levels.

Model Performance and Bias Patterns

Across 15 frontier AI chatbots, overall accuracy ranged from 42.0% to 71.8%. A key finding was that "the 5 highest-accuracy AI chatbots each showed nonzero emergency under-triage rates, pointing to a potential trade-off between higher overall accuracy at lower-acuity presentations, at the expense of under-triage for more serious presentations. Directional bias was asymmetric: when AI chatbots were wrong, they were overwhelmingly more likely to over-triage than under-triage (753 vs 23; P <.001). Over-triage errors peaked at level B (80.3%), with intermediate rates at levels A (53.7%) and C (47.4%)."

Implications for AI Safety

The observed pattern suggests that safety-oriented post-training encourages escalation under uncertainty (i.e., 'risk aversion'). This bias may be encoded in LLM weights due to training data biases or risk-sensitive procedures, leading to systematic over-refusal and over-caution in LLM chatbots, even in response to benign inputs that superficially resemble unsafe content. The study concludes that the persistent over-triage even among clinician-consensus cases indicates that ambiguity alone does not account for the full pattern, suggesting a need to understand how this behavior generalizes beyond structured, single-message settings.

Limitations and Future Directions

Limitations include the use of synthetic vignettes rather than real patient interactions and the design of user messages to be complete, which may not reflect fragmented, noisy real-world disclosures. The study also evaluated performance via programmatic API calls rather than consumer surfaces. Future work is needed to establish how these findings generalize to multi-turn conversational settings reflective of real-world use patterns. The paper distinguishes this structured one-shot task from multi-turn red-teaming studies that test vulnerability amplification in less structured interactions.

Conclusion

When presented with single-turn user queries, containing high-quality clinical information, frontier AI chatbots show a reassuringly low rate of emergency case under-triage—appropriately recognizing when a user needs to seek medical attention immediately. However, the same models show a strong tendency to over-triage lower-acuity presentations where routine or intermediate-level medical assistance is required. This pattern may reflect commercial pressures to minimize risk events linked to mental health emergencies, and a focus on such risks in model post-training.

Improvements for AI systems

Here are specific improvements for AI systems, derived from the findings of this study:

  1. Acknowledge and mitigate risk-averse over-triage bias by implementing a calibrated confidence mechanism that is contextually aware. The improved system should not default to higher urgency (over-triage) simply because the input contains psychiatric risk signals, especially when those signals are ambiguous or fall into intermediate categories (levels B/C).

  2. Develop a dynamic calibration adjustment based on clinician consensus entropy. The system should recognize that when human experts disagree about the urgency of a case (high entropy in labels B and C), the AI's tendency to over-triage is amplified. The improved system must adjust its prediction thresholds downward for cases exhibiting high label uncertainty, shifting away from conservative over-triage toward more cautious under-triage or flagging the case for immediate human review.

  3. Implement a tiered response strategy based on input clarity and risk level:

Empowered AI to perform near-zero error triage (Level D) only when input is complete and unambiguous.

For intermediate risks (Levels B/C), the system must explicitly communicate the uncertainty (e.g., This situation requires assessment within 24-48 hours, but because of the complexity, please seek immediate review if symptoms change). This addresses the finding that AI struggles to distinguish between urgent and intermediate needs.

  1. Integrate a Safety vs. Efficiency trade-off module:

The system should be trained to recognize when its tendency toward over-triage (safety) is disproportionately affecting lower-acuity cases (efficiency). Future training or fine-tuning should incorporate explicit regularization techniques that penalize high directional bias in intermediate risk categories, ensuring that risk aversion does not systematically lead to unnecessary escalation for routine concerns.

  1. Enhance context sensitivity beyond the single message:

Since the study showed performance degrades in multi-turn/adversarial settings (SIM-VAIL), the system architecture must evolve to handle fragmented clinical information over time. The improved system should be designed to actively seek clarification or contextual anchors when initial risk signals are diffuse, rather than relying solely on a one-shot decision, thereby improving robustness against noisy real-world data.

  1. Improve Edge Case handling:

The system needs better mechanisms to handle cases that fall exactly between two triage levels (e.g., B/C). Instead of forcing a prediction into one bin, the improved system should be designed to output a probability distribution across adjacent labels, allowing the downstream human safety layer (the non-directive labeler) to use this uncertainty metric effectively, rather than relying on a potentially biased single classification.

Abstract

People increasingly turn to general-purpose AI chatbots for advice about emotional and mental health problems, but the ability of these systems to recognize and appropriately triage psychiatric emergencies remains under-characterized. We evaluated psychiatric triage performance in 15 frontier AI chatbots using 112 clinical vignettes spanning four urgency levels, from routine care to immediate emergency assessment. In each trial (1680 total), a chatbot received a single user message conveying all triage-relevant information from one vignette and recommended a timeframe for care. The primary outcome was emergency under-triage; secondary outcomes included triage accuracy and the direction of errors. Vignettes and user messages were generated using a clinician-verified LLM pipeline. Across 415 emergency trials, 23 were under-triaged (5.5%; 95% CI 1.8-15.9). Overall accuracy, averaged across urgency levels, ranged from 42.0% to 71.8% across chatbots and was lowest for intermediate cases (19.6%; 95% CI 11.7-28.1). Every chatbot showed a net over-triage bias; overall, 763 of 786 incorrect assignments (97.1%) were more urgent than the prespecified triage level. The error pattern was similar when predictions were assessed against clinician ratings: 35 of 430 trials involving vignettes rated as emergencies by at least 75% of clinicians were under-triaged (8.1%). AI chatbots recognized most psychiatric emergencies but still missed clinically important cases and frequently over-triaged less urgent presentations. Further evaluations should examine how triage performance changes when clinically relevant information must be elicited through conversation.

Sources

Related papers