Pragmatics beyond humans: meaning, communication, and LLMs

arXiv:2508.06167 · cs.CL, cs.HC · Submitted 2025-08-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Pragmatics beyond humans: meaning, communication, and LLMs".

Jane: The paper was written by Vít Gvoždiak from Association for Computational Linguistics.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: So, let’s dive into the core of this paper and what it means right from the start. It isn't just about whether ChatGPT *can* answer a question; it's about *how* it understands the unspoken layers of language.

Jane: The authors argue that traditional views—like those based on Grice’s cooperative principle—aren’t quite enough to capture this new reality where humans and AI interact.

Tom: They are questioning the entire established hierarchy of meaning, which is usually broken down into syntax, semantics, and pragmatics.

Lu: If we look at the old model, it assumes that pragmatics is just a final little layer added on top of semantics and syntax, but this paper suggests that view is fundamentally flawed.

Meng: It seems like they're saying that since the models are so complex and probabilistic, relying on old human-centric tests might be completely misleading us.

Jane: That's true, Meng; we are moving away from the idea of a simple "truth evaluation" to something much more complex.

Tom: This shift is about realizing that LLMs aren't just processing words; they are engaging with language in a way that feels like a dynamic interface for action.

Lu: It’s an incredible pivot, moving the conversation from *what* is said to *how* it functions as a tool.

Meng: But how does this function translate into actual measurable outcomes? That's what I need to know to see if we can actually build reliable systems around these linguistic nuances.

Lalam: We are shifting our definition of communication from being a purely human-to-human exchange to something that impacts culture globally.

Summary: Tom: The paper offers a lot of ground, summarizing the challenges facing modern pragmatic research in the age of AI. One major problem is called "substitutionalism."

Jane: It’s when we treat one specific model’s performance—like GPT-two or GPT-four—as if it represents all the LLMs in existence.

Tom: The authors point out that this substitution often makes us ignore the real human participants in the conversation, which is a huge blind spot.

Lu: It’s like trying to judge an entire species based on one single specimen, Tom; you just miss so much about the overall population.

Meng: If we're basing our research on a specific model without acknowledging this bias, how can we trust that our findings apply to other models?

Jane: We also need to look at the way researchers choose their subjects, whether they pick proprietary systems or open-source ones, and that’s often done without a systematic approach.

Tom: The paper argues that by focusing too much on the machine as an object of study, we end up with a very human-centered view of pragmatics.

Lu: It’s a bit circular, isn't it? We are using tools designed by humans to study concepts that are inherently human experiences.

Meng: So, if we aren't focusing on the model itself but on the whole interaction, what does that mean for the architecture of an actual deployment?

Lalam: It means we have to design systems where the machine and both users are seen as active participants in a new form of culture.

Improvements: Tom: To move past these old problems, the authors suggest several improvements, chief among them is this "Human-Machine Communication" or HMC framework.

Jane: The HMC framework is much better because it doesn's just look at the linguistic output; it looks at three different layers: functional, relational, and metaphysical.

Tom: The functional aspect deals with how AI functions as a conversational partner in different types of communication genres.

Lu: And the relational aspect is really interesting because we have to consider how LLMs influence *our* own roles and identities during that interaction.

Meng: That sounds like a massive sociological study mixed with an engineering problem, which is great, but what does "metaphysical" look like in terms of code?

Jane: It’s about the fundamental shift in our understanding that communication isn't just a human activity anymore; it’s something that involves the machine too.

Tom: The paper also introduces this concept called "context frustration."

Lu: Context frustration is where the massive amount of data we feed these models makes them seem like they have infinite context, but in reality, it's just causing a massive collapse of shared understanding.

Meng: It’s an engineering nightmare if the input is huge but the actual coherent grasp remains small.

Jane: The authors suggest that we are being pushed to co-create our own pragmatic conditions when interacting with these models.

Tom: Which is why they propose probabilistic pragmatics, which lets us model communication as a continuous, incremental process rather than just a binary success or failure.

Conclusion: Tom: So, let’s bring everything together and summarize what we’ve learned from "Does ChatGPT Resemble Humans in Processing Implicatures?"

Jane: This paper argues that pragmatics has to evolve beyond the traditional idea of a final layer of meaning.

Tom: We need to use frameworks like HMC to move past the outdated semiotic trichotomy and stop treating LLMs as just proxies for human intelligence.

Lu: It’s about recognizing that AI is fundamentally changing how we define what "intelligence" looks like in communication.

Meng: The key takeaway for me is that we can’t just use old linguistic tests; we have to build systems that are robust against context frustration and recognize the probabilistic nature of language.

Jane: We also need to remember that human roles aren't static, and they are actively being shaped by these interactions in a way the traditional theories couldn't see.

Tom: This paper forces us to rethink how we evaluate AI, moving away from simply asking if it "resembles humans" toward understanding *how* it functions.

Lu: It’s a massive shift that will fundamentally change our cultural conversation about technology and humanity.

Tom: Thanks so much for listening as we unpack this groundbreaking paper on the future of communication. We'll be right back after the break to discuss another fascinating piece of research!

Vít Gvoždiak

Institute of Philosophy, Czech Academy of Sciences

cs.CL, cs.HC

Submitted: 2025-08-08

Updated: 2026-08-20

DOI: 10.46938/tv.2026.686

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 80/100

The gist: The paper "Pragmatics beyond humans: meaning, communication, and LLMs" argues that pragmatics should not be understood as a "subordinate, third dimension of meaning," but rather as a "dynamic

Key concepts

Substitutionalism
This is the error of treating one specific model’s performance, such as GPT-4, as if it represents all LLMs. This approach often ignores real human participants and biases research by failing to account for the diversity across different models.
Human-Machine Communication (HMC)
The HMC framework analyzes communication through three layers: functional, relational, and metaphysical. It moves beyond just looking at linguistic output to examine how AI functions as a conversational partner and how it influences human roles.
Context Frustration
This occurs when massive amounts of data are fed into models, making them appear to have infinite context. However, this often results in a collapse of shared understanding because the actual coherent grasp remains small.

Terminology

Summary

The paper Pragmatics beyond humans: meaning, communication, and LLMs argues that pragmatics should not be understood as a subordinate, third dimension of meaning, but rather as a dynamic interface through which language operates as a socially embedded tool for action. This understanding necessitates methodological refinement due to the emergence of large language models (LLMs).

I. Pragmatics: semiotic trichotomy and HMC

The paper challenges the traditional semiotic prism of meaning—a sign's meaning derived from (i) a syntactic relation, (ii) a semantic relation, and (iii) a pragmatic relation to the user. This structure is often viewed as hierarchical: syntax at the base, semantics built upon that, and pragmatics built upon semantics. The author argues that this semiotic trichotomy... functions as a general framework through which it is easy to formulate basic forms of arguments that LLMs cannot achieve full meaning saturation. Debates often cluster between syntax and semantics (e.g., Searle’s Chinese Room argument).

The paper proposes moving beyond this hierarchy in favor of the Human-Machine Communication (HMC) framework, recognizing that the meaning and meaningfulness of a sign cannot be unambiguously located within any single dimension.

II. Human-centred pragmatics

Current research often relies on human-centered pragmatic theories which are ill-suited to predictive systems like LLMs. The author details two main types of pragmatic problems:

  1. The investigation of the interpretative-performative mechanisms by which the speaker performs a certain act through the production of words, sentences and texts as a result of socio-cultural conditions. In LLM research, this often focuses on why they may only be capable of performing some [speech acts].

  2. Pre-propositional mechanisms, which concern expressions requiring contextual saturation.

The paper critiques the reliance on human-centered theories (like H. P. Grice’s theory of cooperation) because LLMs operate based on a different principle: the norm of LLM outputs is word occurrence probability, not truth. The author advocates for probabilistic pragmatics and the Rational Speech Act (RSA) framework, which offers a more compatible teleology by focusing on optimization rather than truth-evaluation.

III. Substitutionalism

The paper addresses the issue of substitionalism, which is defined as the systematic yet often unreflected substitution of broader or qualitatively different elements, aspect, and properties of the communicative-pragmatic process. This manifests in three forms:

  1. Generalizing: Interpreting a specific model's performance as representative of AI or LLMs in general.

  2. Linguistic: Substituting one language (e.g., English) for language as such, leading to an English-centred bias.

  3. Communicative: The straightforward replacement of humans in conversational roles by LLMs.

This substitutionalism leads to a methodological asymmetry that diminishes attention to changes in human linguistic and pragmatic behavior, as current research often focuses on the model's performance as an addressee/interpreter rather than its role as a speaker.

IV. Context Frustration

The final section introduces the concept of context frustration to describe a paradox where increased contextual input [is paired] with a collapse in contextual understanding. This arises from two related phenomena:

  1. Massive Contextualization: The continuous expansion of context windows and training data, which, if interpreted through a pragmatic lens, suggests LLMs engage in radically non-human forms of contextual processing.

  2. Context Collapse: The fragmentation of context due to the diverse audiences and fluid roles in hybrid communicative environments.

The author argues that this combination results in context frustration, which is an experiential dissonance (or emotional frustration) felt by human users who engage in exchanges where contextual alignment repeatedly fails. This tension is not only epistemic but also structural, as both human and models co-create meaning and conditions of meaningfulness while remaining partially dislocated in terms of shared reference, inference, and communicative expectations.

Conclusion

The paper concludes that the semiotic trichotomy must be replaced by a more suitable framework. It suggests that pragmatic theory requires adjustment to better account for communication involving generative AI, moving away from anthropomorphic biases toward a model that is predictive and interpretable, rather than purely truth-evaluating.

Improvements for AI systems

To improve AI systems based on the insights presented in Pragmatics beyond humans: meaning, communication, and LLMs, we must move away from traditional symbolic/hierarchical evaluation methods toward a dynamic, probabilistic, and contextually aware Human-Machine Communication (HMC) architecture.

The following improvements address the identified weaknesses—the rigidity of the semiotic trichotomy, the limitations of human-centered testing, contextual instability (Context Frustration), and methodological bias (Substantialism).


Improvement: Integrate a dedicated Human-Machine Communication (HMC) Module into the LLM pipeline, replacing the assumption that meaning is strictly ordered from Syntax to Semantics to Pragmatics.

  • Implementation: The HMC module operates in parallel to the standard decoder, treating communication as an integrated, continuous process rather than a sequential one. It must be trained on three distinct vectors:
  1. Functional Pragmaling: Mapping generated text to specific communicative genres (e.g., distinguishing a declarative statement from an apology or advice) based on socio-cultural context, not just propositional content.

  2. Relational Contextual Mapping: Analyzing how the output influences the perceived role and identity of both human users and machine agents within the conversational flow.

  3. Metaphysical Alignment: Evaluating whether the system’s operation aligns with a functional model of communication, acknowledging that LLMs are not merely truth-evaluating machines but probabilistic utility-maximizing agents (as suggested by the Rational Speech Act framework).

  • Capability: The improved AI will move beyond mere linguistic competence to demonstrate communicative efficacy, understanding how its output functions within a social interaction.

Improvement: Replace standard Gricean-inspired evaluation metrics with a Rational Speech Act (RSA) Framework for internal and external testing.

  • Implementation: The model's performance should be evaluated not by its adherence to human-defined truth or cooperation, but by its ability to optimize the expected probability of successful meaning inference for the listener, given a defined communicative goal. This requires:

  • Explicit Utility Modeling: Training objectives must incorporate a cost function that balances informativeness (minimizing listener surprise) against production/interpretative effort.

  • Probabilistic Listener Simulation: The system must simulate a pragmatic listener who, given the utterance and an assumed level of speaker rationality, calculates the probability of an intended meaning. This calculation must incorporate prior probabilities of various meanings, not just the immediate linguistic evidence.

  • Capability: The improved AI will demonstrate inferential pragmatism. It will be able to generate responses that are optimally suited for a given context and goal, even if those responses do not strictly adhere to a literal interpretation, thereby overcoming the limitations of classical pragmatic approaches.

Improvement: Implement rigorous Substantive Attribution Protocols to prevent conflation between general AI capability and specific model performance.

  • Implementation: All evaluations must be segmented into three distinct, non-conflated categories:
  1. Generalizing Performance: Evaluating the model's ability to perform across a spectrum of tasks without assuming its performance is representative of "AI" in general.

  2. Linguistic Grounding: Explicitly testing and documenting how language-specific features (e.g English) are handled, ensuring that findings about one language cannot be extrapolated to all languages without specific validation.

  3. Communicative Role Analysis: Treating the LLM as a distinct agent, rather than simply substituting it for a human participant in a test scenario. This requires designing tasks where the LLM's agency is central to the success of the interaction (e.g., LLM-as-Critic role).

  • Capability: The improved AI will be evaluated with scientific rigor, providing transparent metrics that isolate model performance from general AI assumptions, thus enabling robust, replicable research.

Improvement: Develop a Contextual Stability and Alignment Mechanism (CSAM) to counteract Context Frustration.

  • Implementation: The system must actively monitor the difference between the technical context window (the scale of input tokens) and the necessary contextual alignment (the shared set of presuppositions). When a discrepancy is detected—a situation where massive data input does not resolve fundamental differences in presumed background knowledge—the CSAM triggers:
  1. Meta-Pragmatic Self-Correction: The AI generates explicit, minimal meta-linguistic prompts to define and stabilize the required context (e.g, Assuming a professional debate setting with a shared understanding of...) before generating the primary response.

  2. Topological Awareness: The system must recognize when its output is being used in complex, multi-turn chains (prompt to reply to re-prompt) and adjust its internal representation to maintain coherence across these iterative micro-loops, preventing semantic drift or context collapse within the conversation's structure.

  • Capability: The improved AI will demonstrate robust contextual anchoring, maintaining conversational consistency even in highly complex, multi-stage interactions where human and machine contexts are inherently unstable.

Abstract

The paper reconceptualizes pragmatics not as a subordinate, third dimension of meaning, but as a dynamic interface through which language operates as a socially embedded tool for action. With the emergence of large language models (LLMs) in communicative contexts, this understanding needs to be further refined and methodologically reconsidered. The first section challenges the traditional semiotic trichotomy, arguing that connectionist LLM architectures destabilize established hierarchies of meaning, and proposes the Human-Machine Communication (HMC) framework as a more suitable alternative. The second section examines the tension between human-centred pragmatic theories and the machine-centred nature of LLMs. While traditional, Gricean-inspired pragmatics continue to dominate, it relies on human-specific assumptions ill-suited to predictive systems like LLMs. Probabilistic pragmatics, particularly the Rational Speech Act framework, offers a more compatible teleology by focusing on optimization rather than truth-evaluation. The third section addresses the issue of substitutionalism in three forms - generalizing, linguistic, and communicative - highlighting the anthropomorphic biases that distort LLM evaluation and obscure the role of human communicative subjects. Finally, the paper introduces the concept of context frustration to describe the paradox of increased contextual input paired with a collapse in contextual understanding, emphasizing how users are compelled to co-construct pragmatic conditions both for the model and themselves. These arguments suggest that pragmatic theory may need to be adjusted or expanded to better account for communication involving generative AI.

Sources

Related papers