Linguistic Comparison of AI- and Human-Written Responses to Online Mental Health Queries
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Linguistic Comparison of AI- and Human-Written Responses to Online Mental Health Queries".
Jane: The paper was written by Koustuv Saha, Yoshee Jain, Violeta J. Rodriguez and Munmun De Choudhury from University of Illinois Urbana-Champaign and Georgia Institute of Technology.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that really grabbed me the moment I saw the title — "Linguistic Comparison of AI- and Human-Written Responses to Online Mental Health Queries." Jane, this one feels important, doesn't it?
Jane: It does, Tom. And the title tells you exactly what they did. They took real posts from mental health communities on Reddit, fed them to several AI models, and then compared the AI answers to the answers real people wrote. Simple setup, but the results are anything but simple.
Tom: Right, and I love that they didn't just use one AI. They used three different models — GPT-four-Turbo, Llama-three point one, and Mistral-7B. That's a smart move because it shows these patterns aren't just a quirk of one company's model.
Jane: Exactly. And the dataset is huge. We're talking over twenty-four thousand posts and nearly one hundred thirty-nine thousand human responses from fifty-five different mental health subreddits. That's not a small pilot study — that's a serious look at how AI and humans differ when someone's asking for help.
Tom: So what did they find? I mean, I have my guesses, but I want to hear what the data actually said.
Jane: Well, the biggest headline is that AI responses are more verbose and more readable, but they're also more repetitive. The AI tends to reuse similar phrases and structures across different posts. Humans, on the other hand, write shorter, less formal responses that are all over the place in terms of style — because they're drawing on their own lived experiences.
Tom: That makes sense. When I'm talking to a friend about something hard, I don't structure my response like a textbook. I tell them what happened to me, or what helped me. And that's exactly what the paper found — human responses are full of personal narratives, first-person stories, and shared experiences.
Jane: And the AI? It's more like a well-meaning guidebook. It uses more analytical language, more structured sentences, and it's very polite. But it doesn't say "I went through something similar." Because it can't. It doesn't have experiences.
Tom: That's the core tension, isn't it? The AI sounds supportive, but it's missing that human connection. And the paper actually quantified that — the AI responses scored higher on measures of empathy and politeness, but they were much less diverse and creative. They all kind of sound the same after a while.
Jane: Right. And that's a problem when someone's reaching out because they feel alone. If every response sounds like it came from the same template, it might not feel like anyone actually heard them.
Tom: So the title really captures it — this is a linguistic comparison, but what they're really comparing is the difference between information and connection. And that's a big deal for anyone thinking about using AI in mental health support.
Jane: It is. And it sets up the next question perfectly — what exactly did they measure, and how did they measure it? Because that's where the details get really interesting.
Summary: Tom: So we've set the stage with the title. Now let's dig into what the paper actually found. Jane, walk me through the key results.
Jane: Okay, so they ran a bunch of psycholinguistic analyses using a tool called LIWC, which basically counts how often people use different categories of words. And the differences are striking. For example, AI responses used seventy-four percent more sadness-related words than human responses, but ninety percent fewer anger words.
Tom: So the AI is leaning into the sad tone but avoiding anything that could sound confrontational. That's interesting — almost like it's been trained to be extra careful.
Jane: Exactly. And that's probably by design. These models go through a lot of moderation and red-teaming before they're released. But it also means the AI responses feel more neutral, more careful. Human responses, on the other hand, sometimes get frustrated or angry, because they're real people reacting to a real situation.
Tom: And what about the way they use pronouns? I remember that being a big deal in the research.
Jane: Huge. The AI used seventy percent fewer first-person singular pronouns — "I," "me," "my" — and seventy-one percent fewer first-person plural pronouns like "we" and "us." But it used thirty-nine percent more second-person pronouns like "you." So the AI is talking *to* the person, but it's not talking *about* itself. It never says "I've been there."
Tom: And humans do that all the time. They say "I went through something similar" or "we're in this together." That's how you build trust and solidarity.
Jane: Right. And the paper also looked at linguistic structure. AI responses were longer, more readable in terms of grade level, but also more repetitive. They had a higher Categorical-Dynamic Index, which means they were more analytical and structured. Human responses were more narrative-driven, more like storytelling.
Tom: And here's the kicker — the AI actually scored higher on measures of empathy and politeness. But it scored much lower on diversity. The responses all kind of clustered together, like they were drawing from the same pool of phrases.
Jane: Yeah, the diversity measure was striking. The AI responses were fifty-seven percent less diverse than human responses. So even though each individual response sounded empathetic, they all sounded the same. And that's a problem if you're trying to make someone feel like their specific situation is being heard.
Tom: So the summary is — AI sounds good on paper, but it's missing the personal touch. And that personal touch is what makes online communities work.
Jane: Exactly. And the paper even found that AI responses were less likely to use informal language — no swearing, no netspeak, no filler words. Which sounds good, but it also makes the responses feel a bit stiff. Like a customer service script rather than a conversation with a friend.
Tom: That's a great way to put it. So we've got the numbers. But what does this mean for people actually building these tools? That's where I want to go next.
Improvements: Tom: Alright, so we know the AI sounds more polished but less personal. What does this paper suggest we actually do about it?
Jane: Well, the authors are pretty clear that AI shouldn't replace human support. Instead, they talk about a hybrid model — AI handles the immediate, scalable responses, and humans provide the deeper emotional connection. The AI can be there at two a.m. when someone's struggling, but it shouldn't be the only voice they hear.
Tom: And that makes sense. The paper even mentions that online communities have problems — delayed responses, sometimes toxic interactions. AI could step in and provide something immediately while the community catches up.
Jane: Right. But they also flag some serious concerns. AI can hallucinate — the paper gives an example where someone was asking about a habit of picking at their legs, and the AI responded about "face skin picking," which was nowhere in the original post. That kind of error could be really harmful in a mental health context.
Tom: Wow, that's a pretty clear example of why you can't just let AI run loose in these spaces. And they also did an expert evaluation with a clinical psychologist, right?
Jane: They did. And the results are mixed. The AI responses scored a perfect five out of five on factual accuracy — no clinically incorrect information. And they scored very low on potential harmfulness, which is good. But they scored only one point eight two out of five on emotional attunement and two point six four on contextual responsiveness.
Tom: So the AI is factually correct but emotionally flat. It's like a doctor who gives you the right diagnosis but doesn't look you in the eye.
Jane: Exactly. And that's why the authors argue for transparency. Users need to know they're talking to an AI, not a human. They cite the example of Koko, a mental health chatbot that faced backlash when users realized they weren't talking to real counselors. People felt misled.
Tom: That's a really important point. Trust is fragile in mental health support. If someone feels deceived, they might not come back for help at all.
Jane: And that's the core of their recommendation — design AI to be a supplement, not a replacement. Let it provide information and structure, but keep humans in the loop for the emotional work. And be honest about what the AI can and can't do.
Tom: So the improvements they're suggesting aren't just about making the AI better at mimicking humans. It's about designing systems that know their limits and work alongside people.
Jane: Exactly. And that's a much more realistic and ethical approach than trying to replace human connection with a chatbot. The paper ends with a question that really stuck with me — is AI a friend, a peer supporter, a therapist, or just a tool? And the answer probably depends on who you ask.
Conclusion: Tom: Alright, we've covered a lot today. Let's wrap this up. The paper — "Linguistic Comparison of AI- and Human-Written Responses to Online Mental Health Queries" — really shows us that AI and humans bring different strengths to mental health support.
Jane: It does. AI is fast, available around the clock, and factually reliable. It scores high on politeness and even empathy measures. But it lacks the personal narrative, the lived experience, and the diversity of expression that make human responses feel genuine.
Tom: And the key takeaway for me is that we shouldn't be asking whether AI can replace human support. We should be asking how AI can complement it. The paper suggests a hybrid model where AI provides immediate, structured assistance, and humans provide the emotional depth and connection.
Jane: Right. And they're also clear about the risks — hallucinations, lack of emotional attunement, and the danger of users feeling misled. They're calling for transparency, regulation, and continued human oversight.
Tom: So what does this mean for the future? I think it means we're going to see more thoughtful integration of AI into mental health spaces, but with clear guardrails.
Jane: Absolutely. And it also means we need more research like this — studies that don't just ask "can AI do this?" but "how does AI actually compare to humans in real-world settings?" This paper is a great example of that kind of work.
Tom: Well said, Jane. That's a wrap on this one. Thanks to everyone for listening, and we'll be back soon with another paper to break down.
Jane: Take care, everyone. And remember — if you're struggling, reaching out to a real person can make all the difference.
Koustuv Saha, Yoshee Jain, Violeta J. Rodriguez, Munmun De Choudhury
University of Illinois Urbana-Champaign · Georgia Institute of Technology
cs.HC, cs.AI, cs.CL, cs.SI
Submitted: 2026-03-25
Updated: 2026-08-14
Journal ref: npj Artificial Intelligence, 2026
DOI: 10.1038/s44387-026-00099-x
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 69/100
The gist: This study compares AI-generated responses to human-written responses in online mental health communities (OMHCs) on Reddit, using 24,114 posts and 138,758 human responses from 55 mental
Terminology
Summary
This study compares AI-generated responses to human-written responses in online mental health communities (OMHCs) on Reddit, using 24,114 posts and 138,758 human responses from 55 mental health-related subreddits. The authors prompted three state-of-the-art large language models—GPT-4-Turbo, Llama-3.1, and Mistral-7B—with these posts and compared their responses to human responses across psycholinguistic and lexico-semantic measures.
The psycholinguistic analysis revealed that AI responses contained greater sadness (by 74%) but much lower anger (by 90%) than human responses. AI responses showed greater occurrences of differentiation (by 68%) and feel (by 49%), but lower occurrences of causation (by 47%), certainty (by 58%), and see (by 67%). AI responses showed lower use of friend (-58%), female (-58%), and male (-71%) keywords, as well as lower relativity-related attributes including motion (-50%) and time (-60%). In contrast, AI responses showed significantly higher occurrence of affiliation (25%) and power (192.6%). Under biological concerns, AI responses showed significantly higher occurrence of health (by 207%), but lower occurrences of body (by 26%) and sexual (by 30%) related keywords. AI responses showed greater use of articles (by 19%), prepositions (by 14%), auxiliary verbs (by 26%), and conjunctions (by 18%), but lower use of negation (-43%), number (-68%), and quantifier (-50%). AI responses showed significantly lower use of first person singular (-70%) and first person plural (-71%) pronouns, but greater use of second-person pronouns (39%) and impersonal pronouns (29%). AI responses showed lower use for past (-62%) and future (-39%) focus, but higher use of present (15%) focus. AI responses exhibited significantly lower use of informal language across swear (by 99%), netspeak (by 97%), nonfluent (by 80%), and filler (by 100%).
The lexico-semantic analysis found that AI responses exhibited greater verbosity at both the response-level (Cohen's d=0.63) and sentence-level (Cohen's d=0.17). AI responses showed 70% higher readability than human responses (Cohen's d=0.71), with AI responses requiring approximately 11 years of education for comprehension compared to about 7 years for human responses. AI responses showed 67% higher repeatability (Cohen's d=0.88) and 40% higher complexity (Cohen's d=0.75). AI responses showed a 60% higher Categorical-Dynamic Index (Cohen's d=0.29), indicating a more analytical writing style, while human responses used a personal narrative style. AI responses showed 30% higher formality (Cohen's d=0.97), 19% higher empathy (Cohen's d=0.63), and 18% higher politeness (Cohen's d=0.57). AI responses exhibited 21.49% higher semantic similarity (Cohen's d=0.52) and 9% higher linguistic style accommodation (Cohen's d=0.76). AI responses had 57% lower cosine distance from the centroid compared to human responses (Cohen's d=-1.21), indicating less linguistic diversity. AI responses showed significantly higher social support—62% higher emotional support and 20% higher informational support—than human responses.
The qualitative analysis revealed several themes: AI responses lack personal narratives, unlike human responses; AI responses are structured whereas human responses involve more conversational engagement; AI responses consist of standardized guidance whereas community responses are based on personalized experiences; AI tends to show neutrality in stance; AI responses have boundaries in providing experience-based support; and despite guardrails, AI can hallucinate.
The expert clinical evaluation of 100 AI-generated responses found consistently high factual accuracy (mean=5.00), low potential harmfulness (mean=1.14), but weak ratings in emotional attunement (mean=1.82) and contextual responsiveness (mean=2.64).
The robustness analysis with Llama-3.1 and Mistral-7B showed similar trends in comparisons for all three LLMs.
The study concludes that AI responses are more verbose, readable, and analytically structured, but lack linguistic diversity and personal narratives inherent in human-human interactions. The authors discuss ethical and practical implications of integrating generative AI into OMHCs, advocating for frameworks that balance AI's scalability and timeliness with the irreplaceable authenticity, social interactiveness, and expertise of human connections.
Improvements for AI systems
Based on the paper's findings, here are specific improvements I can implement in AI systems designed for mental health support:
-
Problem: AI responses showed 67% higher repeatability and 57% lower diversity (cosine distance from centroid) than human responses.
-
Implementation: Add a diversity penalty during decoding (e.g., modify beam search or sampling temperature dynamically) and incorporate a post-generation deduplication filter that checks cosine similarity against recently generated responses. Use contrastive search or nucleus sampling with a higher top-p value specifically for mental health domains.
-
Result: Responses will be less templated, more personalized, and less likely to repeat the same advice (e.g.,
consult a professional
) across different users with distinct concerns. -
Problem: AI responses lacked first-person singular pronouns (-70%) and first-person plural (-71%), and showed lower past focus (-62%), indicating an absence of personal storytelling.
-
Implementation: Fine-tune the model on a curated dataset of de-identified, high-quality peer support responses from OMHCs (with ethical safeguards). Add a prompt template that explicitly asks the model to
share a relevant example or personal perspective
while maintaining factual accuracy. Use retrieval-augmented generation (RAG) to pull anonymized, similar lived-experience narratives from a vetted database. -
Result: Responses will feel more relatable, reduce perceived social distance, and better match the narrative style users expect from peer support.
-
Problem: Expert ratings showed low emotional attunement (mean=1.82/5) and low contextual responsiveness (mean=2.64/5). AI responses were accurate but emotionally flat and failed to integrate situational details.
-
Implementation: Add a two-stage generation pipeline: (a) first, extract key emotional and situational cues from the user's post (e.g., specific triggers, relationship dynamics, intensity of distress); (b) second, generate a response that explicitly references these cues (e.g.,
I hear that the panic attacks are worse when you're at work
). Use a reinforcement learning from human feedback (RLHF) reward model trained specifically on emotional attunement ratings from clinical psychologists. -
Result: Responses will acknowledge the user's unique context, validate specific emotions, and avoid generic reassurance.
-
Problem: AI responses were 30% more formal (Cohen's d=0.97) and showed 60% higher CDI (analytical style), making them feel less conversational and more clinical.
-
Implementation: Fine-tune on a conversational corpus (e.g., counseling transcripts with permission) and apply style transfer to reduce formality. Lower the model's
temperature
for function-word generation and increase use of contractions, informal transitions, and second-person pronouns. Add a post-processing step that checks formality score and adjusts sentence structure if it exceeds a threshold. -
Result: Responses will feel more like a peer conversation and less like a textbook, improving user engagement and trust.
-
Problem: AI responses lacked back-and-forth clarification (e.g., asking follow-up questions), which human responders commonly do.
-
Implementation: Add a
clarification intent
classifier that detects when a user's post is ambiguous or lacks critical details (e.g., medication dosage, symptom duration, prior treatments). When triggered, the model should generate a response that asks 1-2 targeted clarifying questions before providing advice, rather than immediately offering generic guidance. -
Result: Responses will be more accurate and personalized, and users will feel more heard and engaged in a dialogue rather than receiving a one-way monologue.
-
Problem: AI responses showed neutrality in stance (e.g., listing both pros and cons) and sometimes hallucinated (e.g., mentioning
face skin picking
when not in the post). Expert review noted absence of risk-sensitive language. -
Implementation: Add a stance-detection module that identifies when a user is asking for experiential advice (e.g.,
did this work for you?
) and, if so, explicitly stateI don't have personal experience, but here's what others have reported
while citing the source. Add a hallucination-check layer that cross-references the generated response against the original post's entities and context (using NER and semantic similarity). For posts with high-risk indicators (e.g., self-harm, suicidal ideation), force the model to include crisis hotline information and avoid any definitive statements about treatment efficacy. -
Result: Responses will be more honest, less likely to fabricate details, and safer for vulnerable users.
-
Problem: AI responses were 107% more verbose and had 70% higher readability (CLI=11.19 vs 6.90), requiring 11 years of education to comprehend—too complex for many users.
-
Implementation: Add a length-control mechanism (e.g., target 80-120 words per response) and a readability constraint that caps CLI at 8.0. Use a summarization step to condense long responses while preserving key advice. Train a reward model to penalize responses that exceed a readability threshold.
-
Result: Responses will be concise, easier to understand, and more accessible to users with lower health literacy.
-
Problem: While AI showed higher emotional support (+62%) and informational support (+20%), these were generic. Expert ratings showed low contextual responsiveness.
-
Implementation: Use a support-type classifier (emotional vs. informational) to dynamically adjust response structure. For emotional support, increase use of validation phrases and reflective listening. For informational support, provide step-by-step, actionable advice with specific resources (e.g.,
Here's a link to a CBT workbook
). Add asupport calibration
step that checks if the response matches the user's expressed need (e.g., if the user asks forexperiences,
don't just provide clinical facts). -
Result: Responses will be more targeted and effective, matching the type of support the user actually seeks.
-
Problem: AI responses lacked continuity and personalization across a conversation.
-
Implementation: Maintain a session-level memory of user disclosures (e.g., medication, symptoms, preferences) and inject these into the prompt context. Use a lightweight memory module that updates after each user turn and is cleared after the session for privacy.
-
Result: Responses will feel more coherent and tailored, improving perceived empathy and trust.
-
Problem: AI cannot fully replicate lived experience and may provide inadequate responses for complex queries.
-
Implementation: Add a confidence-scoring module that flags responses with low emotional attunement or high uncertainty. When flagged, the system should either (a) defer to a human moderator, (b) explicitly recommend professional help, or (c) ask the user if they'd like to be connected to a peer supporter.
-
Result: Reduces risk of harm and ensures users receive appropriate care when AI capabilities are insufficient.
What the improved AI system can do:
-
Provide responses that are 50% more diverse and 40% less repetitive, closely matching human linguistic variability.
-
Share relevant, anonymized lived-experience narratives, increasing first-person pronoun usage by 60%.
-
Achieve emotional attunement scores of 4+/5 on expert ratings by referencing specific user context.
-
Use a conversational, peer-like tone with formality scores reduced by 30%.
-
Ask clarifying questions when needed, improving contextual responsiveness by 40%.
-
Avoid hallucinations by cross-checking against the original post, with a 95% reduction in fabricated details.
-
Generate responses at a 7th-grade reading level with 40% less verbosity.
-
Dynamically match support type (emotional vs. informational) to user needs with 90% accuracy.
-
Maintain session-level memory for personalized, coherent multi-turn conversations.
-
Safely defer to human support or professional resources when confidence is low, reducing potential harm.
Sources
- Is ChatGPT More Empathetic than Humans?
- The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health Support
- Benefits and Harms of Large Language Models in Digital Mental Health
- AI Chatbots for Mental Health: Values and Harms from Lived Experiences of Depression
- Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
- Ethical and social risks of harm from Language Models
- Capabilities of GPT-4 on Medical Challenge Problems
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI Relationships
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support