Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity

arXiv:2608.13484 · cs.CL, cs.AI · Submitted 2026-08-13 · Read on arXiv

Dananjay Srinivas, Saksham Khatwani, Maria Pacheco

University of Colorado, Boulder

cs.CL, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: The paper "Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity" investigates why large language models (LLMs) fabricate plausible-sounding details about entities

Terminology

Summary

The paper Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity investigates why large language models (LLMs) fabricate plausible-sounding details about entities outside their knowledge boundary instead of retreating to safer, more general claims. The authors frame this failure through a Gricean lens, proposing that a cooperative speaker uncertain about a referent should retreat up the specificity hierarchy, trading informativeness for truthfulness. They ask whether LLMs have the internal ingredients to perform this retreat.

The paper constructs a benchmark dataset from the T-REx partition of LAMA, covering 8 Wikidata relations across 4 domains (people, corporations, products, and skills). For each relation, they generate three levels of contextual grounding for the subject, synthetic substitutions of the subject to simulate entities outside the model's knowledge boundary, and generic substitutions of the object at varying levels of specificity. They verify synthetic entities using the infini-gram API against The Pile dataset, confirming that real entities appear frequently in pretraining data while synthetic entities rarely occur.

The authors probe model activations to answer two questions: (i) do activations encode whether a referent falls inside the knowledge boundary, and (ii) do they anticipate the specificity of the referent about to be generated. Using linear probes (Logistic Regression classifiers) on Pythia models ranging from 70M to 12B parameters, they find that the answer to both questions is yes. Specifically, model activations can strongly predict whether an entity has occurred in pretraining data, with models larger than 2 billion parameters achieving over 90% AUROC. The best predictive layers are roughly just before to the model's middle layer. Additionally, activations can reliably predict whether the model will generate a specific or generic completion, with predictive accuracy rising strongly as layer depth increases, and the 12B model outperforming the 1.4B model.

Despite encoding both signals, the authors find that the two signals are not reconciled in the model's actual generation behavior. Models overwhelmingly prefer specific referents regardless of whether the entity is known or unknown, as measured by both perplexity and an extrinsic surprisal elicitation test. The generation behavior analysis shows that in the real and synthetic cases, the model overwhelmingly prefers to be informative at the cost of truthfulness. For synthetic cases, this means every specific response is likely an undesired outcome. The surprisal test over candidate completions confirms that the model overwhelming prefers specific completions in real and synthetic scenarios, with larger models showing a stronger preference for specific completions while smaller models actually prefer generic ones. This specificity bias persists across all context lengths and increases with longer contexts, possibly indicating overconfidence.

The authors conclude that the substrate for a Gricean retreat is present, but the policy that would act on it is not. They position their findings as a first step toward Gricean alignment, which they define as training or steering objectives that couple knowledge-boundary awareness to referent-specificity during generation. The paper's contributions include constructing a benchmark for testing Gricean retreat behavior, showing that LLM activations encode both knowledge boundary status and upcoming specificity, demonstrating that models produce specific referents regardless of boundary status, and positioning the work as a foundation for Gricean alignment objectives.

Improvements for AI systems

Improvements to AI Systems:

  1. Add a knowledge-boundary gate to generation decoding.
  • Train a lightweight linear probe (as in the paper) on the model’s middle-layer activations to detect whether the current subject entity is inside the pretraining distribution.

  • During decoding, if the probe signals outside boundary, force the decoder to suppress specific object tokens (e.g., named entities, dates, numbers) and instead sample from a restricted set of generic hypernyms (e.g., a company, a person, a product).

  • Resulting capability: The AI will automatically retreat to safe, general statements when asked about obscure or novel entities, reducing hallucination without needing external retrieval.

  1. Implement a specificity-aware contrastive loss during fine-tuning.
  • Use the paper’s benchmark (real vs. synthetic entities with matched relations) to create training pairs: same relation, one known entity, one unknown entity.

  • Add a loss term that penalizes the model when it assigns higher probability to a specific completion for an unknown entity than for a known entity, while rewarding generic completions for unknown entities.

  • Resulting capability: The model learns an explicit policy to trade informativeness for truthfulness, internalizing the Gricean retreat as a learned behavior rather than relying on post-hoc decoding tricks.

  1. Add a specificity calibration module to the output layer.
  • Train a second probe (using the paper’s finding that later layers predict specificity well) to estimate the probability that the next token is specific vs. generic.

  • At inference, compare this estimate to the knowledge-boundary probe output. If the two disagree (e.g., high specificity probability but low boundary confidence), apply a penalty to specific tokens and boost generic alternatives.

  • Resulting capability: The system can dynamically adjust its referent specificity in real time, even for unseen prompts, by reconciling the two internal signals the paper shows are present but unused.

  1. Create a Gricean alignment fine-tuning dataset from the paper’s benchmark.
  • For each synthetic entity, generate a retreat target: the same relation but with the object replaced by a generic hypernym (e.g., What does X do? → X is involved in some activity).

  • Fine-tune the model on these pairs using standard supervised learning, with a curriculum that starts with short contexts and increases context length (since the paper shows specificity bias worsens with longer contexts).

  • Resulting capability: The AI will naturally produce generic completions for unknown entities across all context lengths, directly addressing the observed overconfidence.

  1. Add a knowledge-boundary watermark to the model’s internal state.
  • Use the probe’s AUROC (over 90% for >2B models) to create a binary flag in the hidden state at the middle layer.

  • Propagate this flag through the transformer’s residual stream to the final layer, where it biases the softmax toward generic tokens when the flag is unknown.

  • Resulting capability: The model’s generation becomes explicitly conditioned on its own knowledge boundary, making the retreat behavior robust to prompt phrasing and context length variations.

What the improved AI system can do:

  • Answer questions about obscure or novel entities with safe, generic statements (e.g., I’m not familiar with that specific person, but they are likely a professional in some field) instead of fabricating plausible names, dates, or affiliations.

  • Maintain truthfulness even under long, detailed prompts that currently trigger overconfident specificity.

  • Provide a tunable truthfulness vs. informativeness knob (via the probe thresholds) that developers can adjust per application—e.g., strict retreat for medical or legal queries, lenient for creative writing.

  • Automatically flag its own uncertainty in its output (e.g., by using hedging phrases) when the knowledge-boundary probe indicates low confidence, improving user trust.

  • Serve as a foundation for downstream tasks like fact-checking, summarization, and dialogue, where hallucination on rare entities is a known failure mode.

Abstract

When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to safer, more general claims. We frame this failure through a Gricean lens: a cooperative speaker who is uncertain about a referent retreats up the specificity hierarchy, trading informativeness for truthfulness. We ask whether LLMs have the ingredients to perform this retreat. Using a T-REx-based benchmark that varies entity familiarity and referent specificity, we probe models to answer two questions: (i) do their activations encode whether a referent falls inside the knowledge boundary, and (ii) do they anticipate the specificity of the referent they are about to generate? We find that the answer to both is yes, but the two signals are not reconciled in generation. Models overwhelmingly prefer specific referents even when the entity is unknown to them, and do so even when offered correct generic alternatives. The substrate for a Gricean retreat is present, but the policy that would act on it is not. We position our findings as a first step toward Gricean alignment, training or steering objectives that couple knowledge-boundary awareness to referent-specificity during generation.

Sources

Related papers