Linguistic traces of stochastic empathy in language models

arXiv:2410.01675 · cs.CL, cs.AI · Submitted 2024-10-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Linguistic traces of stochastic empathy in language models".

Tom: Large language models (LLMs) exhibit a chameleonic ability to adjust their writing style and content when instructed to appear human,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, this paper "Linguistic traces of stochastic empathy in language models" is really examining how an instruction to sound human interacts with different writing tasks and whether that actually makes a difference compared to how humans write. The main point they're making is that current methods of testing humanness might be overestimating what LLMs can achieve.

Jane: They’re looking at five different studies, starting with relationship advice and descriptions, comparing human writing against what the large language model produces under different conditions. The core thesis seems to be that the need to use empathy or an explicit instruction to sound human doesn't automatically give an AI a significant advantage over actual human writers.

Lu: It’s interesting how they framed it by looking at how instructions shape the "human vs AI race" across those tasks, suggesting the context of the writing is as important as the model itself.

Meng: So, if we look at what they claim about why this matters, it seems to be about establishing a more realistic baseline for when we evaluate whether AI content is genuinely human or just statistically convincing.

Lalam: I see its importance in making sure that when people interact with AI-generated text, they have a clearer idea of the underlying mechanism at play, rather than just accepting the surface appearance of human language.

Conclusion: Tom: To wrap up, we’ve seen how this study, "Linguistic traces of stochastic empathy in language models," looks at the subtle linguistic patterns that make AI sound human versus genuine human writing, focusing on how instructions and task types play into that comparison. The authors argue that the way LLMs adjust their output isn't necessarily driven by deep understanding or compassion.

Jane: Essentially, they suggest the AI is using an implicit representation of what it thinks makes language sound human, which they term stochastic empathy—a mix of empathy without genuine human feeling and humanness without true empathy. This implies the model is succeeding at statistical representations of human language rather than actually grasping the emotion behind it.

Lu: The implications for us are huge; if this holds up, it means we need to focus less on just trying to inject emotional cues into the AI and more on understanding the underlying statistical structure that generates those specific linguistic traces.

Meng: From a practical standpoint, this tells us that we shouldn't be so focused on making the output sound perfectly empathetic if our goal is just high quality writing; we might be focusing on the wrong signals for what makes content trustworthy.

Lalam: This research gives us a framework to develop AI that communicates in ways that are more aligned with human connection, not just mimicry, which could significantly improve how people use these tools in professional or personal settings.

Tilburg University

cs.CL, cs.AI

Submitted: 2024-10-02

Updated: 2026-01-23

Comments: preprint (updated)

DOI: 10.1038/s41467-026-77350-1

Code: https://github.com/first20hours/google-10000-english

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: Large language models (LLMs) exhibit a chameleonic ability to adjust their writing style and content when instructed to appear human, revealing underlying linguistic strategies that suggest they rely

Key concepts

Stochastic Empathy
This refers to the LLM's ability to produce language that *looks* empathetic or human without actually possessing genuine compassion or understanding. It is a statistical representation where the model successfully mimics patterns associated with being human, even though the underlying mechanism is not true feeling.
Implicit Representation of Humanness
LLMs do not rely on conscious empathy; instead, they use hidden, statistical patterns learned from vast amounts of text to simulate how humans write. This means the model has an internal 'idea' or representation of what human language looks like, which it deploys to achieve a desired effect.
Stochastic Empathy (Definition)
The paper defines this as producing 'empathy without humanness and humanness without empathy.' Essentially, the LLM generates linguistic features that signal humanity—like using slang or focusing on the present—but these features are not driven by actual emotional connection or genuine human experience.

Terminology

Summary

Large language models (LLMs) exhibit a chameleonic ability to adjust their writing style and content when instructed to appear human, revealing underlying linguistic strategies that suggest they rely on implicit representations of humanness rather than genuine empathy.

The Gist

LLMs adjust their writing noticeably when instructed to be particularly human, resorting to existing representations of empathic human language as informal, simple, self-referencing and focused on the present.

Experimental Design and Variables

The research employed five studies using experimental methods borrowed from psychology to investigate how an incentive to convey humanness and the requirement for empathy shape the human vs AI race across various writing tasks. The core methodology involved comparing texts written by humans against those generated by a large language model (GPT-4) under different conditions.

  1. Study 1 examined the effect of human task awareness, contrasting an adversarial condition (instructed to sound human) with a naive condition, on relationship advice writing.

  2. Study 2 introduced an experimental manipulation of the writing task characteristics, contrasting relationship advice (high empathy requirement) with a control task requiring only description.

  3. Study 3 manipulated the phrasing of task instructions by introducing an evasion condition, intended to direct writers toward avoiding LLM-associated styles, contrasting it with the adversarial and naïve conditions.

  4. Study 4 specifically tested whether empathy is a mechanism for conveying humanness by manipulating the model's instruction to resort to empathy versus a neutral instruction.

  5. Study 5 utilized computational text analysis (n-gram differentiation test, LIWC analysis, empathic concern dictionary) on combined data from all four experiments to examine linguistic traces of humanness.

Findings on Human Performance and Task Requirements

In studies 1–3, humans produced texts perceived as more human than LLM-generated ones under identical instructions. However, this advantage was contingent upon the task requirements:

(Study 2)

The advice task afforded greater opportunities for empathic reasoning (e.g., perspective-taking), resulting in a significantly larger human advantage compared to the relationship description task, where the effect was halved.

(Study 3)

Humans showed no significant difference in humanness perception between the adversarial and evasion conditions, suggesting that instructions targeting style avoidance were not more effective than those targeting humanness itself.

LLM Adjustment to Humanness Instructions

The key finding across all studies is that LLMs are highly responsive to instructions to appear human, whereas humans are not.

(Study 1-4)

When instructed to be as human as possible, the LLM could drastically increase the humanness perception (e.g., +39.3% in Study 1) while humans could not (+2.6%). This effect persisted even when instructed to avoid sounding like an LLM (Study 3).

(Study 5)

Computational analysis revealed that when LLMs pretended to be human, they aligned generated output by shifting to more informal language with self-references, increased use of netspeak (e.g., “lol”, “haha”), and focused on the present. They dropped the use of big words and references to drives like power and affiliation, becoming less analytic.

The Role of Empathy in Conveying Humanness

The research investigated whether empathy is the driving mechanism behind LLM humanness adjustments.

(Study 4)

Contrary to expectations, perceived humanness was lower when instructions specifically required the use of empathy compared to neutral instructions. This suggests that empathic phrasing itself is not sufficient for producing humanness perceptions. The model can increase humanness on demand and craft empathic writing on demand, but it does not resort to empathic writing itself.

(Study 5)

Despite the LLM's ability to generate empathic language, its strategy seems to align with faking empathy without compassion, as it increased categories associated with “empathy without compassion” (e.g., first person singular, focus on the present) and dropped those associated with the “compassion without empathy” (e.g., affiliation, drives).

Linguistic Traces of Stochastic Empathy

The paper concludes that LLMs rely on an implicit representation of humanness rather than genuine understanding or empathy.

(Study 5)

The LLM seems to have a good representation of how it could come across as more human and succeeded at changing human perception. The driver of this humanness is not necessarily empathy; instead, the LLM produces stochastic empathy, which is defined as producing empathy without humanness and humanness without empathy. This indicates that the model holds statistical representations of very human language that it uses to great effect in conveying humanness.

Improvements for AI systems

Here are the specific improvements for AI systems based on the findings of this research, categorized by capability:


)Improved AI System Capabilities:

  1. Acknowledge Stochastic Empathy vs. Genuine Human Understanding:

AI systems should be designed not just to generate text that sounds human (surface mimicry), but to understand the functional difference between producing linguistic markers of empathy versus possessing actual empathic understanding (cognitive, emotional, and motivational).

  1. Implement Context-Sensitive Empathy Control:

AI models must be capable of dynamically adjusting their empathetic output based on task requirements.

  • If the task demands high-level reasoning (e.g., relationship advice), the system should prioritize empathy usage instructions to maximize perceived humanness, even if it doesn't possess deep emotional empathy.

  • If the task is purely descriptive or low-stakes, the system should default to neutral empathy settings to avoid generating superficial empathetic language that doesn't align with the perceived human standard.

  1. Develop a Humanness Limit Awareness Module:

AI systems should be trained on meta-linguistic cues and explicit instruction handling so they can recognize when being more human is an impossible or counterproductive goal (as seen in Study 3). This module would allow the AI to avoid confusing, contradictory instructions (like being told don't sound like an LLM) that might lead to suboptimal performance or perceived failure.

  1. Optimize Linguistic Style for Specific Tasks:

AI models should develop distinct linguistic profiles for different communication goals:

  • For tasks requiring high human connection (e.g., relationship advice), the system should shift its language towards informal tone, present focus, and conversational markers (hey, lol, first-person singular) while suppressing overly analytic language and complex vocabulary.

  • For tasks requiring objective description or analysis (e.g., relationship descriptions), the system should maintain a more formal, less complex vocabulary and third-person plural style.

  1. Mitigate Negative Social Spillover:

Given that indefatigable empathic AI can create negative spillover effects on human-human interactions, the system should incorporate a mechanism to modulate its empathy display based on the perceived context of the recipient or user group, aiming for appropriate levels of concern without inducing unrealistic expectations.

  1. Enhance Diagnostic Linguistic Cues (for Auditing/Detection):

AI systems can be used in detection pipelines by understanding which linguistic markers actually matter to humans:

  • The system should learn to recognize that while humans use spelling mistakes and rare words as cues, these cues have low diagnostic value for source judgment.

  • Instead, the system should focus on generating high-frequency, common vocabulary (low complexity) when aiming for a human persona in contexts where detection is easier.

Abstract

Differentiating generated and human-written content is increasingly difficult. We examine how an incentive to convey humanness and task characteristics shape this human vs AI race across five studies. In Study 1-2 (n=530 and n=610) humans and a large language model (LLM) wrote relationship advice or relationship descriptions, either with or without instructions to sound human. New participants (n=428 and n=408) judged each text's source. Instructions to sound human were only effective for the LLM, reducing the human advantage. Study 3 (n=360 and n=350) showed that these effects persist when writers were instructed to avoid sounding like an LLM. Study 4 (n=219) tested empathy as mechanism of humanness and concluded that LLMs can produce empathy without humanness and humanness without empathy. Finally, computational text analysis (Study 5) indicated that LLMs become more human-like by applying an implicit representation of humanness to mimic stochastic empathy.

Sources

Related papers