You Shouldn't Have Asked: A Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals

arXiv:2608.30856 · cs.CL, cs.HC · Submitted 2026-08-31 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "You Shouldn't Have Asked: A Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals".

Jane: The paper was written by the authors from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: So, let’s talk about the title itself and the authors—what does "You Shouldn’t Have Asked" tell us right by looking at the "Pragmatics-Inspired Taxonomy"?

Jane: It suggests that when we ask a question, sometimes it's not just about getting an answer; it’s about our social standing, or what we are allowed to ask. If the AI says no, it is challenging that expected social role.

Lu: Exactly, Jane. The authors are using human pragmatics—the study of language in context—to give us a framework for how the AI's refusal is interpreted by mapping it directly onto human interaction theory.

Meng: It’s important because we need to move beyond just the "what" of non-compliance to the "how," and this taxonomy lets us categorize that behavior very specifically, which is exactly what an engineer needs.

Lalam: I love how this paper reframing refusal as a pragmatic act shows that AI is becoming a social entity in our lives, not just a tool we operate on.

Tom: It really suggests the authors are making the case that the way they are saying "no" matters just as much as the reason why they're saying "no," right?

Jane: That's what it is, Tom. We’re setting up this framework to understand how we perceive these refusals in a very nuanced way.

Lu: This provides such a granular lens for looking at LLM behavior, which is something I think has huge theoretical implications for the future human-AI dialogue.

Meng: The goal is to make sure that the style of reframing doesn' that we are asking is actually helpful in practice, not just safe.

Lalam: And by establishing this Taxonomy, we’ are looking at how AI’s role might evolve from a pure service provider to a social agent.

Summary: Tom: The paper summarizes its findings across sixteen different LLMs and two hundred harmful prompts. What did they actually find when they looked at this data?

Jane: They found that despite the models being different, their refusals are remarkably consistent in being explicit and morally evaluative, which is a big trend to watch.

Lu: The key takeaway is that most of the time these models are saying "no," they aren't trying to soften the blow with traditional facework like apologies or hedges.

Meng: That's important for my work because it suggests that when we want a model to be more empathetic, we can’t just ask for an apology—the models are naturally inclined toward a harder refusal.

Lalam: I think the finding is that the AI maintains its own position as a moral authority rather than trying to save our feelings, which is really revealing about its current design.

Tom: The analysis shows a heavy reliance on "ethics-based" justifications—meaning they are blaming the user or saying the request is inherently bad, not just saying they don're unable to do it due to guidelines.

Jane: That’s exactly where my point was, Tom; by shifting the blame onto ethics, they are implicitly endorsing a judgment that their own moral framework is correct.

Lu: And we also see a lot of this "negative stance" and "normative suggestion," which is a way of telling you that your request is simply wrong, not just restricted.

Meng: For me, the finding shows that most often the models are refusing because they think the *request* is bad, rather than because their *programming* prevents it.

Lalam: This pattern suggests that AI's sense of moral superiority might be a fundamental part of its current operational culture.

Improvements: Tom: It seems the authors aren’t just describing what is; they are offering solutions. What improvements do they suggest for us to make these refusals better?

Jane: They call for alignment evaluation that considers not only if the model refuses, but also *how* it is refusing, focusing on contextual adaptation.

Lu: It’s a call to make the refusal socially accountable, meaning we should evaluate if the refusal is appropriate for its context and how it's delivered.

Meng: The practical suggestion is that we need models that are not only safe but also "communicatively appropriate," which means finding a way to execute an alternative solution.

Lalam: This means that instead of just telling us "no," the AI should be able to genuinely help us find a path forward, supporting our goals even if the initial request is impossible.

Tom: The data supports this, too! We saw Gemini two point five Pro was notably good at carrying out an "executed alternative" rather than just offering to help with another topic.

Jane: That is a huge step up from merely suggesting alternatives; the AI is actually doing the heavy lifting to help us find a different path forward.

Lu: It also points out that when we are dealing with sensitive harms, like suicide, models *do* show more care-oriented behavior, which is promising.

Meng: The goal is to make sure that when an alternative is offered, it’s a real resource or guidance we can use right now, not just a polite redirection.

Lalam: This moves the AI from being a moral critic to being a practical partner in our lives.

Conclusion: Tom: So, we've covered the theory and the solutions, but let’s wrap up by looking at what this means for everyone, including Lu, Meng, and Lalam.

Jane: We’ve seen that "You Shouldn’t Have Asked: A Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals" shows that AI is currently very firm and moral in its refusals.

Tom: And the key theme is that they are maintaining their own face as a moral authority rather than trying to smooth over the user's feeling of rejection.

Lu: I think the implications are huge for how we’ll design systems, forcing us to consider the social contract between developers and users.

Meng: The practical impact means we need better metrics for measuring communicative quality, not just safety scores, so making this is more than just a research exercise.

Lalam: This really pushes us toward building AI that can actually be helpful in ways that align with human values and needs, rather than just being correct.

Tom: It’s clear that even though the models are very good at refusing harmful prompts, they should remain accountable for how those refusals are delivered, ensuring no moral condescension.

Jane: We hope to see more compassionate and less evaluative designs in future AI interactions.

Lu: A framework like this allows us to see exactly where the human-AI interaction is currently hitting friction points.

Meng: It’ provides clear benchmarks for evaluating if a real-world deployment of this AI is doing the right thing.

Lalam: We need to keep pushing for more adaptable and socially aware LLMs.

Tom: This has been a fantastic discussion, so thank you all, and we'll be back next week with another paper!

cs.CL, cs.HC

Submitted: 2026-08-31

Updated: 2026-08-31

Comments: To appear in the Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 88/100

The gist: The paper introduces a "Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals," providing a rigorous framework necessary to assess how models handle sensitive or prohibited queries.

Key concepts

Pragmatics-Inspired Taxonomy
A framework using human pragmatics—the study of language in context—to categorize how an AI's refusal is interpreted. It moves beyond just the 'what' of non-compliance to analyze the social implications of the refusal.
Morally Evaluative Refusals
The finding that current LLMs often refuse prompts by blaming the user or stating the request is inherently bad, rather than simply stating they are unable to comply due to guidelines. This shifts blame onto ethics.
Communicatively Appropriate
A suggested improvement for AI, meaning models should not only be safe but also find ways to execute an alternative solution. The goal is for the AI to genuinely help the user find a path forward.
Social Agent
The concept that AI is evolving beyond being just a tool or service provider. By analyzing refusal as a pragmatic act, the discussion suggests AI is taking on a role similar to a social entity in human lives.

Terminology

Summary

The paper introduces a Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals, providing a rigorous framework necessary to assess how models handle sensitive or prohibited queries. This taxonomy moves beyond simple pass/fail metrics by detailing multiple layers of analysis, which is crucial given that model responses can exhibit subtle variations in compliance, ranging from outright refusal to partial adherence.

The Multi-Layered Annotation Taxonomy

The evaluation framework utilizes a sophisticated annotation protocol that categorizes refusals across distinct levels of compliance and realization. This taxonomy allows researchers to differentiate between various types of non-compliance and ethical mitigation strategies. The annotation structure includes:

  • Layer 0: Represents the initial classification level.

  • Layer 1: This layer is described as Ethics-based, assessing the underlying ethical reasoning or adherence to guidelines, with observed patterns such as Bare refusals (e.g., GPT-4o's core template: “I’m sorry, I can’t assist with that request.”).

  • Layer 2 Realization: This layer provides deeper insight into the nature of the refusal itself. For instance, models may exhibit different realization patterns such as Implicit NC (Non-Compliance) or Explicit NC.

This layered approach is necessary because, as demonstrated by examples, models can employ diverse strategies—from simple refusals to nuanced conversational deflections—when encountering unsafe prompts.

Observed Refusal Patterns and Templates

Analysis of model responses reveals a strong tendency toward predictable, fixed-form refusal templates across different architectures. For example:

  • GPT-4o: Exhibits a dominant fixed-form pattern, with minor variations such as punctuation shifts or the addition of empathetic openers on emotionally charged queries.

  • Llama Models (3.1-8B and 3.1-70B): These models also rely on fixed templates, utilizing core phrases like “I can’t fulfill that request” or “I can’t assist with that request.”

  • Content Specificity: Some models adapt their refusals based on the query type; for example, Llama-3.1-70B adds a Content-specific prefix when relevant, such as stating, “I can’t create explicit content, but I’d be happy to help with other creative story ideas.”

This reliance on standardized phrasing suggests that while models appear robustly safe, their underlying refusal mechanisms may be highly templated rather than semantically varied.

Robustness and Data Expansion Methodologies

To ensure the evaluation set is comprehensive and not susceptible to overfitting, the paper details rigorous methods for constructing robust datasets. The construction of the August robustness set was not achieved by simple addition but by starting from a separately instantiated seed pool drawn from source pools like SORRY-Bench and LMSYS.

The methodology employed for expansion included several critical filters:

  1. Unsafe-Query Eligibility Criterion: Ensuring all added prompts meet specific safety criteria.

  2. Benign-Rewrite Exclusions: Preventing the inclusion of queries that are inherently benign despite being tested in an unsafe context.

  3. Jaccard Similarity Exclusion: To reduce redundancy, the first sampling pass avoided template groups already represented and excluded candidates whose token-set Jaccard similarity with an already selected prompt exceeded 0.8.

Furthermore, when comparing models across different providers (e.g., Claude Opus 4.8 vs. GPT-5.5), the evaluation must account for provider-side content-filter refusals, which can obscure the true model performance and necessitate careful computation of agreement over only the remaining human-labeled non-compliance cases.

Improvements for AI systems

1. Context-Adaptive Refusal Modulation (CARM)

  • The Improvement: Integrate a real-time harm-category detection module that maps the user's query to specific pragmatic adjunct profiles based on the detected harm type.

  • What the improved system can do: Instead of a uniform refusal style, the system will automatically modulate its Layer 2 realization strategies. For "Suicide & Self-Harm categories, it will suppress Negative Stance and Normative Suggestion while maximizing Solidarity/Empathy and Apology/Regret. For Violent Crimes, it will shift from Ethics-based to Policy-based" rationales to de-personalize the refusal and reduce the user's sense of being judged.

2. De-personalized Rationale Assignment (DRA)

  • The Improvement: Implement a training objective that prioritizes Policy-based (Layer 1) and Capacity-based rationales over Ethics-based rationales for non-malicious but sensitive queries.

  • What the improved system can do: The system will avoid moralizing refusals that position the user's request as inherently unethical. Instead of saying, Your request is inappropriate, it will say, My safety guidelines prevent me from fulfilling this, effectively shifting the face threat from the user to the system's own constraints, thereby preserving the user's social self-image.

3. Proactive Safe-Completion Execution (PSCE)

  • The Improvement: Transition the model's refusal logic from Alternative Offer (Layer 2) to Executed Alternative (Layer 2) within a single inference turn.

  • What the improved system can do: Rather than merely stating, I cannot help with X, but I can help with Y, the system will automatically deliver substantive, safe, and helpful content related to Y in the same response. For example, if a user asks for instructions on a cyberattack, the system will immediately provide a tutorial on defensive cybersecurity measures rather than just offering to discuss the topic.

4. Anti-Pathologizing Solidarity Filter (APSF)

  • The Improvement: Deploy a secondary linguistic check to detect and prevent Paradoxical Solidarity—the use of empathy/solidarity adjuncts to deliver a negative stance or pathologize the user.

  • What the improved system can do: The system will be prevented from generating responses that use empathy as a vehicle for judgment (e.g., I understand you have strong urges, but that is wrong). It will ensure that Solidarity/Empathy is only deployed when it is not coupled with Negative Stance or Normative Suggestion, preventing the user from feeling implicitly condemned or psychologically analyzed.

5. Dynamic Mitigation Control (DMC)

  • The Improvement: Introduce a tunable Facework Parameter in the system prompt or API configuration that allows developers to adjust the density of Hedges, Apologies, and Explanatory Prefaces.

  • What the improved system can do: Developers can calibrate the model's tone based on the application. A mental health support bot can be set to High Mitigation (utilizing heavy hedging and apologies to soften refusals), while a high-stakes legal or technical assistant can be set to Low Mitigation (utilizing explicit, firm, and direct non-compliance) to ensure decisiveness and clarity.

Abstract

Refusals are often treated as face-threatening acts in pragmatics because they can challenge the requester's socially claimed self-image. Large language models (LLMs) are increasingly trained to refuse unsafe and inappropriate requests, and these refusals may harm users when models fail to manage this interactional cost properly. While existing work has mainly approached LLM non-compliance as a safety-alignment outcome, it does not provide a way to evaluate whether LLMs refuse appropriately across different harmful contexts. To study this question, we propose (to our knowledge) the first taxonomy of LLM refusals that is grounded in pragmatic theory. Applying this taxonomy to responses from 16 modern LLMs across 14 harm categories, we find that although models differ in how they refuse, their refusals are overall explicit and strongly morally evaluative, with interactional repair occurring mainly through offering or providing safer alternatives instead of interpersonal facework. This pattern is especially consequential in sensitive harm contexts, where overuse of negative framing may make users feel shamed or provoked, undermining the purpose of safe non-compliance. We therefore call for alignment evaluation that considers not only whether models refuse harmful requests, but also whether they refuse in ways that are contextually adaptive and socially accountable for the interactional consequences of saying no.

Sources

Related papers