Speech-based Psychological Crisis Assessment using LLMs
summary
The gist
This paper proposes a novel large language model (LLM)-based framework for automated psychological crisis level classification using authentic speech data from support hotline calls.
In short
A novel framework uses a large language model (LLM) to classify psychological crisis levels from support hotline calls. It combines speech transcription with injected paralinguistic emotional cues derived from a SpeechLLM. This approach provides objective, consistent assessments by explicitly translating non-verbal emotional signals into text for the LLM to analyze.
Key concepts
- Paralinguistic Injection Strategy
- This strategy transforms raw speech transcripts by adding explicit descriptions of non-verbal cues like 'Trembling voice' or 'obvious sobbing.' These cues are guided by a triage assessment form to make the speaker's affect clear and directly linked to their words, reducing ambiguity for the model.
- Reasoning-Enhanced Training Strategy
- The model is trained not just on predicting the correct crisis level but also on generating step-by-step clinical rationales. This auxiliary task forces the LLM to justify its prediction by referencing both the injected paralinguistic transcript and established TAF criteria across affective, behavioral, and cognitive domains.
- Data Augmentation
- To handle limited authentic data (154 calls), segments are broken into fixed-duration chunks. This creates over 900 speech chunks. Predictions for subjects are then aggregated using majority voting to ensure the final subject-level assessment is robust and temporally continuous.
- Multimodal Data Preprocessing
- This initial step involves two stages: first, converting raw audio into text using an Automatic Speech Recognition (ASR) system. Second, a SpeechLLM injects emotional descriptions into these texts to create a 'paralinguistically enriched transcription' before the LLM classification begins.
Terminology used across episodes
This episode discusses
- Speech-based Psychological Crisis Assessment using LLMs · Paper Radio
- An Exploratory Deep Learning Approach for Predicting Subsequent Suicidal Acts in Chinese Psychological Support Hotlines
- Deep Learning and Large Language Models for Audio and Text Analysis in Predicting Suicidal Acts in Chinese Psychological Support Hotlines
- Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines
- Cognitive-Mental-LLM: Evaluating Reasoning in Large Language Models for Mental Health Prediction via Online Text
- Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
- Step-Audio-R1 Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- Qwen2.5-Omni Technical Report
- LoRA: Low-Rank Adaptation of Large Language Models
The paper
Speech-based Psychological Crisis Assessment using LLMs · Read on arXiv
Terumi Chiba, *, Yang Luo, *, Ziyun Cui, *, Yongsheng Tong, **, Chao Zhang, **
Tsinghua University · Peking University Huilongguan Clinical Medical School · WHO Collaborating Centre for Research and Training in Suicide Prevention
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Speech-based Psychological Crisis Assessment using LLMs".
Jane: This paper proposes a novel large language model (LLM)-based framework for automated psychological crisis level classification using authentic speech data from support hotline calls.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the title and who wrote this paper, "Speech-based Psychological Crisis Assessment using LLMs." It really shows exactly what they are trying to do—using speech analysis combined with a large language model to figure out if someone is in a crisis. Jane The authors are from Tsinghua University and Peking University's Clinical Medical School, which gives them a solid foundation in the medical side of things, which I think is crucial for this kind of work.
Lu: Their background in clinical settings probably helps them understand what actual distress looks like when it comes across a voice, which is something purely technical models often miss Lu.
Meng: So, how does this title translate into something practical for the people actually using these hotlines? What's the immediate impact we should be looking at?
Lalam: I think the real implication here is moving away from subjective human judgment toward a more consistent, objective way of triage. If we can get that consistency, it really helps standardize care across different centers Lalam.
The paper's summary: Tom: So, summarizing the core idea of "Speech-based Psychological Crisis Assessment using LLMs," they propose a framework where raw audio is first transcribed into text, and then they inject paralinguistic descriptions—like sobbing or trembling—into those transcripts before an LLM predicts the crisis level. Jane It's like taking the spoken words and adding the non-verbal emotional context directly into what the AI reads, which should give it a much richer picture of what’s going on.
Lu: That’s smart because it bridges that gap between just reading text and actually understanding how a person sounds when they are distressed, which is where a lot of the nuance lives Lu.
Meng: But I wonder about the transcription step first; getting an accurate text transcript from audio is one hurdle, and then injecting those specific emotional cues needs to be very precise so it doesn't just add noise.
Lalam: The paper mentions that this approach aims to provide objective evaluations essential for efficient triage and resource planning in mental health crisis services Lalam. That focus on efficiency through objectivity is what really caught my eye.
The paper's improvements: Tom: They suggest a couple of big improvements, starting with the paralinguistic injection strategy, where they guide those emotional cues using the triage assessment form to make sure the injected descriptions actually match clinical criteria. Jane That sounds like they are trying to make sure the model isn't just guessing emotions but is basing them on established psychological frameworks.
Lu: And then they introduce a reasoning-enhanced training strategy, which is really clever; instead of just asking for a classification, they train the model to generate step-by-step clinical rationales conditioned on both the enriched transcript and those criteria Lu.
Meng: So, they aren't just building a predictor; they’re forcing the AI to show its work by generating a diagnostic reasoning chain, which should make their output much more trustworthy for clinicians.
Lalam: It also mentions using data augmentation by dividing the audio into non-overlapping chunks of fixed duration to create more training material, which helps them deal with having relatively limited authentic clinical data Lalam.
Conclusion: Tom: So, wrapping up on "Speech-based Psychological Crisis Assessment using LLMs," the main point is that explicitly translating those vital non-verbal cues into text effectively bridges the gap between what we hear and what an LLM can understand, yielding a macro F1-score of zero point eight zero two. Jane It confirms that combining semantics with affect in this way allows the text LLM to perform better at crisis assessment than models looking only at one modality alone.
Lu: I think the real power is in how it handles that modality gap; by forcing the model to process both the words and those vocal characteristics together, it gets a much more holistic view of a caller's state Lu.
Meng: From my side, seeing that performance compared to baselines like the Zero-shot LLM Baseline shows that this added complexity actually translates into significant practical gains in accuracy for high-stakes classification tasks Meng.
Lalam: I see this leading to systems where support resources can be deployed much more precisely because the assessment is more consistent and robust, which is a huge step forward for service quality Lalam.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization