Linguistic Comparison of AI- and Human-Written Responses to Online Mental Health Queries
summary
The gist
This study compares AI-generated responses to human-written responses in online mental health communities (OMHCs) on Reddit, using 24,114 posts and 138,758 human responses from 55 mental
This episode discusses
- Linguistic Comparison of AI- and Human-Written Responses to Online Mental Health Queries · Paper Radio
- Is ChatGPT More Empathetic than Humans?
- The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health Support
- Benefits and Harms of Large Language Models in Digital Mental Health
- AI Chatbots for Mental Health: Values and Harms from Lived Experiences of Depression
- Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
- Ethical and social risks of harm from Language Models
- Capabilities of GPT-4 on Medical Challenge Problems
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI Relationships
The paper
Linguistic Comparison of AI- and Human-Written Responses to Online Mental Health Queries · Read on arXiv
Koustuv Saha, Yoshee Jain, Violeta J. Rodriguez, Munmun De Choudhury
University of Illinois Urbana-Champaign · Georgia Institute of Technology
DOI: 10.1038/s44387-026-00099-x
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Linguistic Comparison of AI- and Human-Written Responses to Online Mental Health Queries".
Jane: The paper was written by Koustuv Saha, Yoshee Jain, Violeta J. Rodriguez and Munmun De Choudhury from University of Illinois Urbana-Champaign and Georgia Institute of Technology.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that really grabbed me the moment I saw the title — "Linguistic Comparison of AI- and Human-Written Responses to Online Mental Health Queries." Jane, this one feels important, doesn't it?
Jane: It does, Tom. And the title tells you exactly what they did. They took real posts from mental health communities on Reddit, fed them to several AI models, and then compared the AI answers to the answers real people wrote. Simple setup, but the results are anything but simple.
Tom: Right, and I love that they didn't just use one AI. They used three different models — GPT-four-Turbo, Llama-three point one, and Mistral-7B. That's a smart move because it shows these patterns aren't just a quirk of one company's model.
Jane: Exactly. And the dataset is huge. We're talking over twenty-four thousand posts and nearly one hundred thirty-nine thousand human responses from fifty-five different mental health subreddits. That's not a small pilot study — that's a serious look at how AI and humans differ when someone's asking for help.
Tom: So what did they find? I mean, I have my guesses, but I want to hear what the data actually said.
Jane: Well, the biggest headline is that AI responses are more verbose and more readable, but they're also more repetitive. The AI tends to reuse similar phrases and structures across different posts. Humans, on the other hand, write shorter, less formal responses that are all over the place in terms of style — because they're drawing on their own lived experiences.
Tom: That makes sense. When I'm talking to a friend about something hard, I don't structure my response like a textbook. I tell them what happened to me, or what helped me. And that's exactly what the paper found — human responses are full of personal narratives, first-person stories, and shared experiences.
Jane: And the AI? It's more like a well-meaning guidebook. It uses more analytical language, more structured sentences, and it's very polite. But it doesn't say "I went through something similar." Because it can't. It doesn't have experiences.
Tom: That's the core tension, isn't it? The AI sounds supportive, but it's missing that human connection. And the paper actually quantified that — the AI responses scored higher on measures of empathy and politeness, but they were much less diverse and creative. They all kind of sound the same after a while.
Jane: Right. And that's a problem when someone's reaching out because they feel alone. If every response sounds like it came from the same template, it might not feel like anyone actually heard them.
Tom: So the title really captures it — this is a linguistic comparison, but what they're really comparing is the difference between information and connection. And that's a big deal for anyone thinking about using AI in mental health support.
Jane: It is. And it sets up the next question perfectly — what exactly did they measure, and how did they measure it? Because that's where the details get really interesting.
Summary: Tom: So we've set the stage with the title. Now let's dig into what the paper actually found. Jane, walk me through the key results.
Jane: Okay, so they ran a bunch of psycholinguistic analyses using a tool called LIWC, which basically counts how often people use different categories of words. And the differences are striking. For example, AI responses used seventy-four percent more sadness-related words than human responses, but ninety percent fewer anger words.
Tom: So the AI is leaning into the sad tone but avoiding anything that could sound confrontational. That's interesting — almost like it's been trained to be extra careful.
Jane: Exactly. And that's probably by design. These models go through a lot of moderation and red-teaming before they're released. But it also means the AI responses feel more neutral, more careful. Human responses, on the other hand, sometimes get frustrated or angry, because they're real people reacting to a real situation.
Tom: And what about the way they use pronouns? I remember that being a big deal in the research.
Jane: Huge. The AI used seventy percent fewer first-person singular pronouns — "I," "me," "my" — and seventy-one percent fewer first-person plural pronouns like "we" and "us." But it used thirty-nine percent more second-person pronouns like "you." So the AI is talking *to* the person, but it's not talking *about* itself. It never says "I've been there."
Tom: And humans do that all the time. They say "I went through something similar" or "we're in this together." That's how you build trust and solidarity.
Jane: Right. And the paper also looked at linguistic structure. AI responses were longer, more readable in terms of grade level, but also more repetitive. They had a higher Categorical-Dynamic Index, which means they were more analytical and structured. Human responses were more narrative-driven, more like storytelling.
Tom: And here's the kicker — the AI actually scored higher on measures of empathy and politeness. But it scored much lower on diversity. The responses all kind of clustered together, like they were drawing from the same pool of phrases.
Jane: Yeah, the diversity measure was striking. The AI responses were fifty-seven percent less diverse than human responses. So even though each individual response sounded empathetic, they all sounded the same. And that's a problem if you're trying to make someone feel like their specific situation is being heard.
Tom: So the summary is — AI sounds good on paper, but it's missing the personal touch. And that personal touch is what makes online communities work.
Jane: Exactly. And the paper even found that AI responses were less likely to use informal language — no swearing, no netspeak, no filler words. Which sounds good, but it also makes the responses feel a bit stiff. Like a customer service script rather than a conversation with a friend.
Tom: That's a great way to put it. So we've got the numbers. But what does this mean for people actually building these tools? That's where I want to go next.
Improvements: Tom: Alright, so we know the AI sounds more polished but less personal. What does this paper suggest we actually do about it?
Jane: Well, the authors are pretty clear that AI shouldn't replace human support. Instead, they talk about a hybrid model — AI handles the immediate, scalable responses, and humans provide the deeper emotional connection. The AI can be there at two a.m. when someone's struggling, but it shouldn't be the only voice they hear.
Tom: And that makes sense. The paper even mentions that online communities have problems — delayed responses, sometimes toxic interactions. AI could step in and provide something immediately while the community catches up.
Jane: Right. But they also flag some serious concerns. AI can hallucinate — the paper gives an example where someone was asking about a habit of picking at their legs, and the AI responded about "face skin picking," which was nowhere in the original post. That kind of error could be really harmful in a mental health context.
Tom: Wow, that's a pretty clear example of why you can't just let AI run loose in these spaces. And they also did an expert evaluation with a clinical psychologist, right?
Jane: They did. And the results are mixed. The AI responses scored a perfect five out of five on factual accuracy — no clinically incorrect information. And they scored very low on potential harmfulness, which is good. But they scored only one point eight two out of five on emotional attunement and two point six four on contextual responsiveness.
Tom: So the AI is factually correct but emotionally flat. It's like a doctor who gives you the right diagnosis but doesn't look you in the eye.
Jane: Exactly. And that's why the authors argue for transparency. Users need to know they're talking to an AI, not a human. They cite the example of Koko, a mental health chatbot that faced backlash when users realized they weren't talking to real counselors. People felt misled.
Tom: That's a really important point. Trust is fragile in mental health support. If someone feels deceived, they might not come back for help at all.
Jane: And that's the core of their recommendation — design AI to be a supplement, not a replacement. Let it provide information and structure, but keep humans in the loop for the emotional work. And be honest about what the AI can and can't do.
Tom: So the improvements they're suggesting aren't just about making the AI better at mimicking humans. It's about designing systems that know their limits and work alongside people.
Jane: Exactly. And that's a much more realistic and ethical approach than trying to replace human connection with a chatbot. The paper ends with a question that really stuck with me — is AI a friend, a peer supporter, a therapist, or just a tool? And the answer probably depends on who you ask.
Conclusion: Tom: Alright, we've covered a lot today. Let's wrap this up. The paper — "Linguistic Comparison of AI- and Human-Written Responses to Online Mental Health Queries" — really shows us that AI and humans bring different strengths to mental health support.
Jane: It does. AI is fast, available around the clock, and factually reliable. It scores high on politeness and even empathy measures. But it lacks the personal narrative, the lived experience, and the diversity of expression that make human responses feel genuine.
Tom: And the key takeaway for me is that we shouldn't be asking whether AI can replace human support. We should be asking how AI can complement it. The paper suggests a hybrid model where AI provides immediate, structured assistance, and humans provide the emotional depth and connection.
Jane: Right. And they're also clear about the risks — hallucinations, lack of emotional attunement, and the danger of users feeling misled. They're calling for transparency, regulation, and continued human oversight.
Tom: So what does this mean for the future? I think it means we're going to see more thoughtful integration of AI into mental health spaces, but with clear guardrails.
Jane: Absolutely. And it also means we need more research like this — studies that don't just ask "can AI do this?" but "how does AI actually compare to humans in real-world settings?" This paper is a great example of that kind of work.
Tom: Well said, Jane. That's a wrap on this one. Thanks to everyone for listening, and we'll be back soon with another paper to break down.
Jane: Take care, everyone. And remember — if you're struggling, reaching out to a real person can make all the difference.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization