Towards Explainable Multimodal Depression Recognition for Clinical Interviews

arXiv:2501.16106 · cs.CL · Submitted 2026-08-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Towards Explainable Multimodal Depression Recognition for Clinical Interviews".

Jane: The paper was written by Wenjie Zheng, Qiming Xie, Zengzhi Wang, Jianfei Yu and Rui Xia from Nanjing University of Science and Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We're starting things off with a heavy hitter today called "Towards Explainable Multimodal Depression Recognition for Clinical Interviews." It's a recent paper from Wenjie Zheng and a team of researchers over at Nanjing University of Science and Technology.

Jane: I've been looking at the title, Tom, and it really highlights a gap in how we use technology in medicine. Most of these systems just give you a final answer, but this one wants to show its work.

Tom: Right, Jane, and that's a huge deal for doctors who can't just take a computer's word for it.

Lu: I'm particularly drawn to that "multimodal" aspect mentioned in the title. It means the system isn't just reading a transcript; it's actually processing the audio signals and the visual facial cues all at once.

Jane: That sounds incredibly complex to pull off in a real interview.

Lu: It is, but when you combine the tone of a voice with the way someone's eyes move, you get a much richer picture of their mental state.

Meng: I have to wonder about the practical side of this, though. If we're talking about clinical interviews, how do we ensure these multimodal inputs don't just create a mess of data that's impossible to parse?

Tom: That's a fair point, Meng, and I think that's why they added "explainable" right there in the name.

Jane: Exactly, because if the AI can explain that it noticed a specific change in speech patterns, a clinician can actually verify that observation.

Meng: If it can point to specific evidence, then it becomes a tool for the doctor rather than a replacement for them.

Lalam: This shift toward transparency is what will allow AI to truly integrate into our social fabric. We can't have a culture of blind trust in machines, especially when we're discussing something as sensitive as mental health.

Tom: It's a massive responsibility, and this paper seems to be tackling it head-on.

Jane: We should look closer at how they actually structure this "explanation" they're promising.

Summary: Tom: Moving deeper into "Towards Explainable Multimodal Depression Recognition for Clinical Interviews," the researchers aren't just tweaking old models; they're proposing an entirely new task called EMDRC.

Jane: To make sense of that, you have to understand the PHQ-eight which is a standard eight-item scale doctors use to check for symptoms like sleep issues or changes in appetite.

Tom: So the AI is essentially following the same playbook as a human professional.

Jane: Precisely, and they've built a new dataset called DAIC-Explain to make this happen.

Lu: What's brilliant about the DAIC-Explain dataset is that it doesn't just label someone as "depressed" or "not depressed." It provides a structured summary that covers four specific areas.

Meng: I saw that in the paper, Lu. It includes a history of depression, a symptom assessment, the underlying causes, and even potential action plans.

Tom: That's a lot more detail than a simple classification score.

Meng: It's also a massive amount of work for the researchers. They had to manually annotate these summaries to ensure the AI had a high-quality ground truth to learn from.

Jane: It's a heavy lift, but it's what gives the model its "brain," so to speak.

Lu: I love how they look for the "why" behind the symptoms. If a participant says they can't sleep because of work stress, the model is designed to capture that connection.

Lalam: That's the part that touches the human experience. By identifying that a lack of sleep is tied to career challenges, the AI is actually mapping the context of a person's life.

Tom: It moves the conversation from "what is happening" to "why is this happening."

Jane: And that leads us directly into the specific methods they used to train these models to think that way.

Improvements: Tom: We're getting into the technical heart of "Towards Explainable Multimodal Depression Recognition for Clinical Interviews" now, specifically their PhqMML framework.

Jane: Think of PhqMML as a student who is studying for two different tests at the same time. Instead of just learning to predict depression severity, the model is also learning to classify every single sentence based on those PHQ-eight items.

Lu: That multi-task learning is so clever because it forces the model to pay attention to the specific semantics of each utterance. It's using a Cross-Modal Transformer to weave the text, audio, and vision together into one cohesive understanding.

Tom: And the results they're seeing are pretty staggering.

Meng: I was looking at those numbers, Tom. They achieved a thirteen point seven eight percent absolute increase in the F1 score compared to the previous state-of-the-art model, HiQuE.

Jane: That's a huge jump for a single research paper.

Meng: It really is, and I'm curious if that performance holds up when you don't have a massive training set.

Tom: That's where their second method, PhqCoT, comes in. It's a "training-free" approach that uses Chain of Thought prompting.

Lu: I love the PhqCoT idea. It basically tells a Large Language Model to slow down and reason through each of the eight PHQ items one by one, extracting evidence from the dialogue as it goes.

Jane: It's like asking a student to show every step of their math problem instead of just writing down the answer.

Lalam: This reasoning process is vital. When the model uses a Chain of Thought to justify its score, it creates a bridge of logic that humans can follow.

Tom: It's a dual approach: one for deep, specialized training and one for flexible, reasoning-based inference.

Jane: We've seen how they built it and how it performs, so let's wrap this all up.

Conclusion: Tom: We've reached the end of our look at "Towards Explainable Multimodal Depression Recognition for Clinical Interviews."

Jane: This paper really sets a new standard for what we should expect from AI in the medical field.

Tom: It's not enough to be right; the model has to be able to explain its reasoning to the people who matter most.

Lu: I see this opening up so many doors for specialized AI that can adapt to different mental health contexts and even different languages.

Meng: From an engineering standpoint, the way they've structured the multi-task learning gives us a clear blueprint for building more robust, interpretable systems.

Lalam: Ultimately, this research moves us closer to a future where technology doesn't just process data, but actually respects the complexity of the human condition.

Tom: Thanks for joining us, everyone. We'll see you next time for another deep dive into the latest research.

Jane: Goodbye for now!

Wenjie Zheng, Qiming Xie, Zengzhi Wang, Jianfei Yu, Rui Xia

Nanjing University of Science and Technology

cs.CL

Submitted: 2026-08-20

Updated: 2026-08-21

Code: https://github.com/NUSTM/EMDRC

Importance score: 68/100

The gist: The paper addresses the limitation that "existing MDRC [multimodal depression recognition for clinical interviews] studies mainly focus on improving task performance...

Key concepts

Multimodal Depression Recognition
A system that processes multiple data types simultaneously, such as text transcripts, audio signals, and visual facial cues. By combining voice tone with eye movements and speech, the AI gains a richer understanding of a patient's mental state than text alone could provide.
PHQ-eight
A standard eight-item scale used by doctors to check for depression symptoms, such as sleep issues or changes in appetite. The researchers use this scale to guide the AI, ensuring it follows a professional playbook when assessing patients and providing explanations.
PhqMML Framework
A multi-task learning framework that uses a Cross-Modal Transformer to weave text, audio, and vision together. It learns to both predict depression severity and classify individual sentences based on PHQ-eight items, significantly improving accuracy over previous models.
PhqCoT
A training-free approach that uses Chain of Thought prompting with Large Language Models. It requires the model to reason through each PHQ item one by one, extracting specific evidence from the dialogue to create a logical, human-followable justification for its scores.

Terminology

Summary

The paper addresses the limitation that "existing MDRC [multimodal depression recognition for clinical interviews] studies mainly focus on improving task performance... However, for clinical applications, model transparency is critical, and previous works ignore the interpretability of decision-making processes. To address this, the authors propose an Explainable Multimodal Depression Recognition for Clinical Interviews (EMDRC) task, which aims to provide evidence for depression recognition by summarizing symptoms and uncovering underlying causes. The goal of the EMDRC task is to structured summarize participant’s symptoms based on the eight-item Patient Health Questionnaire depression scale (PHQ-8), and predict their depression severity. The proposed structured summary includes screening for a history of depression, assessment of the scale-listed symptoms, assessment of possible underlying causes, and action plans."

To support this new task, the researchers construct an EMDRC dataset named DAIC-Explain based on the existing MDRC dataset DAIC-WOZ. The authors propose two distinct approaches to tackle the EMDRC task:

  1. PhqMML (Training-based setting): This is a PHQ-aware multimodal multi-task learning framework that captures the utterance-level symptom-related semantic information to help generate dialogue-level summary. The framework leverages advanced LLMs to annotate PHQ-8 item label for each utterance in the dialogue and introduces an auxiliary PHQ-8 item classification task to capture symptom-related semantic information for each utterance. It then utilizes PHQ-aware symptom summaries with acoustic and visual features via a Cross-Modal Transformer (CMT) for dialogue-level depression recognition.

  2. PhqCoT (Training-free setting): This is a PHQ-guided Chain of Thought (CoT) prompting method designed to enhance LLMs’ ability in predicting depression severity while generating high-quality symptom summaries to improve explainability. The method "first guides LLMs to evaluate each PHQ-8 item by assigning scores based on evidence extracted from the dialogue transcript. It then generates symptom summary and determines the severity of depression based on the total score across all items."

Experimental results demonstrate the superiority of our proposed methods over baseline systems on the EMDRC task. Specifically, "PhqMML outperforms previous state-of-the-art MDRC method by 13.78% absolute percentage points on F1 score, demonstrating that incorporating symptom summary not only enhances model performance but also improves interpretability. The study further observes a positive correlation between symptom summary performance and severity prediction results, confirming that symptom summary can indeed help depression severity prediction and improve model interpretability."

Improvements for AI systems

Improvement 1: Transition from Black-Box Classification to PHQ-Aware Multi-Task Learning (PhqMML) Architecture.

  • Technical Implementation: Integrate an auxiliary utterance-level classification head into the encoder to predict specific PHQ-8 item labels (e.g., Sleep Disorder, Appetite Changes) for every segment of dialogue. Use these labels to guide a LongT5-based decoder for structured symptom summary generation, while simultaneously utilizing a Cross-Modal Transformer (CMT) to fuse these symptom-aware textual embeddings with acoustic (Wav2Vec2) and visual (3D facial landmark) features.

  • Improved AI Capability: The system will no longer provide a single, uninterpretable severity score. Instead, it will generate a comprehensive, structured clinical report that includes: (1) a history of depression screening, (2) a detailed assessment of the eight PHQ-8 symptoms, (3) an identification of underlying psychosocial causes (e.g., work stress, financial instability), and (4) evidence-based clinical action plans.

Improvement 2: Symptom-Semantic Cross-Modal Fusion.

  • Technical Implementation: Replace generic multimodal fusion with a Cross-Modal Transformer (CMT) that specifically aligns PHQ-8-labeled semantic textual representations with intra-modal acoustic and visual features. This involves using Self-Attention Transformers (SAT) to model intra-modal interactions before performing cross-modal alignment between the symptom-heavy text and the physiological cues.

  • Improved AI Capability: The system will be able to perform symptom-specific multimodal verification. For example, it can cross-reference a participant's verbal report of psychomotor agitation with specific high-frequency facial micro-movements and acoustic jitter, significantly increasing the precision of the diagnosis by ensuring physiological cues match the reported symptoms.

Improvement 3: PHQ-Guided Chain-of-Thought (PhqCoT) Inference Logic.

  • Technical Implementation: For training-free or zero-shot deployment, implement a sequential reasoning prompting strategy that forces the LLM to evaluate each of the eight PHQ-8 items individually. The prompt must mandate that the model extracts specific dialogue evidence to justify a score (0–3) for each item before synthesizing the final summary and severity prediction.

  • Improved AI Capability: The AI acts as a transparent clinical auditor. In real-time clinical settings, a doctor can review the model's step-by-step reasoning chain to see exactly which part of the conversation led to a Severe Sleep Disorder rating, allowing for immediate verification or correction of the model's logic, thereby mitigating the risk of AI hallucinations in psychiatric diagnosis.

Sources

Related papers