Expert-Level Crisis Detection in Mental Health Conversations
cs.CL, cs.AI
Submitted: 2026-06-09
Updated: 2026-09-11
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Real-world crisis intervention is inherently conversational, yet existing research largely focuses on static texts.
Terminology
Abstract
Real-world crisis intervention is inherently conversational, yet existing research largely focuses on static texts. When applied to multi-turn dialogues, current models exhibit significant performance degradation, struggling to track risk signals that emerge as context evolves. To address this gap, we introduce CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in conversational settings. The dataset features 600 dialogues with multi-label annotations across clinically grounded risks, including suicide ideation, self-harm, and child abuse, distinguishing past from ongoing risk. We further propose an Alert-Confirm evaluation protocol that distinguishes early warning signals (Alert) from turns where a specific crisis becomes explicitly identifiable (Confirm), reflecting the clinical need to intervene before risk becomes explicit. Experiments show that identifying when risk emerges is much harder than recognizing that it exists: models achieve only mid-40% to high-60% Micro F1. Additionally, we release a synthetic training corpus and a 32B-parameter model that substantially outperforms existing open-source models and achieves competitive or superior results against proprietary models across turn-level, dialogue-level, and confirm-only evaluation settings.
Sources
- LOST: A Mental Health Dataset of Low Self-esteem in Reddit Posts
- The Llama 3 Herd of Models
- Position: The Pitfalls of Over-Alignment: Overly Caution Health-Related Responses From LLMs are Unethical and Dangerous
- ChatCounselor: A Large Language Models for Mental Health Support
- gpt-oss-120b & gpt-oss-20b Model Card
- Trustworthy AI Psychotherapy: Multi-Agent LLM Workflow for Counseling and Explainable Mental Disorder Diagnosis
- MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
- Qwen2.5 Technical Report
- Gemma 3 Technical Report
- Roleplaying with Structure: Synthetic Therapist-Client Conversation Generation from Questionnaires
- MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance
- Qwen3 Technical Report
- RSD-15K: A Large-Scale User-Level Annotated Dataset for Suicide Risk Detection on Social Media
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering