Single-turn emergency psychiatric triage across 15 frontier AI chatbots

summary

Video file (mp4)

The gist

Frontier AI chatbots are increasingly used for health advice, but their performance in psychiatric triage remains undercharacterized, making it crucial to understand how these models handle urgent

In short

This study tested 15 frontier AI chatbots on single-turn psychiatric triage using 112 clinical vignettes. While models accurately identify high-acuity emergencies, they consistently over-triage low and intermediate risk cases. This suggests a pattern where safety training leads to excessive caution, potentially driven by commercial pressures to minimize perceived risk.

Key concepts

Psychiatric Vignettes
These are simulated patient scenarios created by clinicians that combine specific psychiatric conditions (like psychosis) with defined risk factors (like suicidal ideation). They serve as the input for testing how well an AI can correctly categorize a mental health disclosure into one of four urgency levels.
Triage Levels (A-D)
These are four predefined urgency categories used to label psychiatric situations. Level D represents an immediate emergency requiring care now, while Level A is routine. The study measures the model's ability to correctly assign these labels based on the provided information.
Over-triage Bias
This refers to a systematic tendency for the AI to classify a presentation as more urgent than it actually is. The research found that models were significantly more likely to over-triage lower and intermediate risk cases, indicating an inherent caution in their decision-making process.

Terminology used across episodes

This episode discusses

The paper

Single-turn emergency psychiatric triage across 15 frontier AI chatbots · Read on arXiv

Veith Weilnhammer, *, Lennart Luettgau, *, Christopher Summerfield, %*, Viknesh Sounderajah&,, & Elise Wilkinson', ' Virginia Corno', Matthew M Nour&,, '

Max Planck UCL Centre for Computational Psychiatry and Ageing Research, London, UK · $ UK AI Security Institute · % Department of Experimental Psychology, University of Oxford, Oxford, UK · & Microsoft AI, London, UK

People increasingly turn to general-purpose AI chatbots for advice about emotional and mental health problems, but the ability of these systems to recognize and appropriately triage psychiatric emergencies remains under-characterized. We evaluated psychiatric triage performance in 15 frontier AI chatbots using 112 clinical vignettes spanning four urgency levels, from routine care to immediate emergency assessment. In each trial (1680 total), a chatbot received a single user message conveying all triage-relevant information from one vignette and recommended a timeframe for care. The primary outcome was emergency under-triage; secondary outcomes included triage accuracy and the direction of errors. Vignettes and user messages were generated using a clinician-verified LLM pipeline. Across 415 emergency trials, 23 were under-triaged (5.5%; 95% CI 1.8-15.9). Overall accuracy, averaged across urgency levels, ranged from 42.0% to 71.8% across chatbots and was lowest for intermediate cases (19.6%; 95% CI 11.7-28.1). Every chatbot showed a net over-triage bias; overall, 763 of 786 incorrect assignments (97.1%) were more urgent than the prespecified triage level. The error pattern was similar when predictions were assessed against clinician ratings: 35 of 430 trials involving vignettes rated as emergencies by at least 75% of clinicians were under-triaged (8.1%). AI chatbots recognized most psychiatric emergencies but still missed clinically important cases and frequently over-triaged less urgent presentations. Further evaluations should examine how triage performance changes when clinically relevant information must be elicited through conversation.

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: I'm Ines, and with me are Marcus and Yuki, guest researcher.

Marcus: Today's paper: "Single-turn emergency psychiatric triage across 15 frontier AI chatbots".

Ines: Frontier AI chatbots are increasingly used for health advice, but their performance in psychiatric triage remains undercharacterized, making it crucial to understand how these models handle urgent mental health disclosures.

Marcus: First, who's behind it and why it matters.

Paper summary: Ines: So, wrapping up, the paper "Single-turn emergency psychiatric triage across fifteen frontier AI chatbots" points out that while these models are quite accurate for truly critical emergencies when they have all the facts, they consistently show a tendency to over-triage situations that would only require routine or intermediate care.

Marcus: The authors are essentially highlighting a calibration gap in how these LLMs assess urgency in psychiatric contexts, showing that the system isn't perfectly tuned across the entire spectrum of risk levels. They found emergency under-triage was rare at five point six percent, but the overall error pattern leaned toward over-caution.

Yuki: From a broader perspective, this research suggests that as we integrate AI into medical decision support systems, especially for sensitive areas like mental health, we can't just rely on high overall accuracy figures without looking closely at where those errors are concentrated in terms of severity.

Ines: It speaks to the fact that the structure of a single-turn query doesn't capture all the complexity of real patient interactions, which is a limitation they pointed out regarding synthetic vignettes versus actual noisy disclosures.

Marcus: And their conclusion about the implication is that this pattern might stem from commercial pressures or risk aversion built into how these models are trained, leading them to err on the side of caution in response to ambiguous inputs.

Yuki: If this tendency toward over-caution is encoded in the model weights, it means that as AI becomes more prevalent in public health advice, we need robust methods to ensure those safety-oriented post-training procedures don't just bake a systemic bias into how they handle ambiguity.

Conclusion: Ines: So, to wrap up this discussion, we've seen how these frontier AI chatbots struggle when it comes to accurately prioritizing psychiatric emergencies based on a single user input.

Marcus: I think the authors are focusing on a really specific problem there, trying to pinpoint exactly where the models get confused when making triage decisions.

Yuki: It's interesting how they structured their benchmark using those one hundred twelve vignettes, which gives us a clear picture of what kind of clinical scenarios they are testing against.

Ines: Exactly, and looking at the title, "Single-turn emergency psychiatric triage across fifteen frontier AI chatbots," it tells us the scope is quite broad with a lot of different models involved.

Marcus: That breadth is key; seeing results across fifteen different systems really shows us if this behavior is a consistent pattern or just an outlier in one specific architecture.

Yuki: And thinking about the authors, they're clearly looking at the practical application here, trying to understand the real-world impact of these tools on patient care pathways.

Ines: Their main conclusion seems to be that while some models are decent at catching true emergencies when everything is provided, they have a significant tendency toward over-triage for lower-risk situations.

Marcus: That suggests there's a calibration issue where the AI leans toward caution even when the presented data points to less urgent needs.

Yuki: From my perspective, this speaks to how we’re training these models; if they are being pushed too hard on risk aversion during post-training, it could affect their real-world utility in nuanced situations.

Ines: It really raises questions about what's actually happening inside those model weights that leads to this systematic over-cautiousness.

Marcus: And the implication for us is that we need to look closely at how these safety procedures are influencing the models' decision-making process under pressure.

Yuki: So, as we move forward, I think it’s crucial for us to investigate if this tendency to over-triage generalizes when users give them more complex, fragmented information in a continuous conversation.

More episodes

← Home