Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content, Evaluated Against Trained Verifiers
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content, Evaluated Against Trained Verifiers".
Jane: Retrieval-augmented generation (RAG) systems often suffer from hallucination, and this research introduces Grounding-Aware Sensitivity by Perturbation (GASP),
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, focusing on that title again, "Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content, Evaluated Against Trained Verifiers," what I'm trying to break down for the listeners is that this research moves the detection from an answer level to a sentence level. That means we can finally identify exactly which claim is unsupported within a longer AI response.
Jane: Exactly, Tom; it’s like moving from checking if an entire essay is good or bad to checking every single paragraph for specific errors. And the fact that it's training-free is important because it doesn't require us to manually label thousands of examples of what constitutes a hallucination.
Lu: The authors are presenting GASP, which they frame as a method that scores each answer sentence by how strongly its likelihood depends on the retrieved evidence, which they call grounding sensitivity <ref:2607.04223#pg0>. This is a very specific way of measuring dependence that goes deeper than simple text statistics.
Meng: The authors mention they use perturbation—they test how the model's likelihood drops when specific context chunks are removed—to measure this sensitivity, which sounds like a very direct way to probe the model’s reliance on particular pieces of information.
Lalam: It’s interesting that they compare their method against trained verifiers; it shows that while we have those separate models, GASP offers a way to get that level of detail without needing to maintain and version another large verification pipeline constantly.
The paper's summary: Tom: So, summarizing the core idea of "Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content, Evaluated Against Trained Verifiers," the paper introduces GASP. Essentially, it calculates grounding sensitivity for every sentence by measuring how much that sentence’s likelihood changes when you either use the full context or when you remove specific chunks of that context.
Jane: To put that in simpler terms for our audience, imagine the AI generates a long answer; this method checks each sentence and asks, "How much did this sentence rely on this specific piece of evidence we found?" If removing that evidence causes a huge drop in the AI’s confidence in that sentence, it suggests the claim is unsupported.
Lu: They define grounding sensitivity using two main features: first, the gap between log-likelihoods when comparing the full context versus no context, which they call gap(y), and second, a local measure called drop(y), which tells us how much confidence drops when just one specific chunk is taken out <ref:2607.04223#pg0>.
Meng: That localization aspect with drop(y) is what I find most practical; it doesn't just tell us the sentence is suspicious, it points directly to the exact chunk of retrieved evidence that was crucial for its generation.
Lalam: It’s a powerful concept because they frame this as measuring how the dynamical system’s attractor reacts when we perturb the evidence that defines it, which gives us a theoretical backbone to this practical measurement.
The paper's improvements: Tom: Now for the improvements they suggest; beyond just creating a detector, they propose using these four grounding-sensitivity features—gap(y), jsd∅(y), drop(y), and jsdloo(y)—to create a final score that is a decreasing function of them. Higher scores mean greater suspicion of hallucination, which is what we want for an audit tool.
Jane: The training-free default detector, GASP-threshold, standardizes these four features and then applies a threshold to their negated sum; this means no one needs to label data for the primary detection mechanism to work right out of the box.
Lu: They also present a supervised variant called GASP-trained that uses a gradient-boosted decision tree classifier to map those four feature vectors directly onto a hallucination score, which adds another layer of refinement if we have some training data available <ref:2607.04223#pg1>.
Meng: The real improvement for us is the explainability part; they show that for any flagged span, you can return the unit that caused the maximum drop in log-likelihood, which acts as a candidate supporting passage instead of just saying "this is wrong."
Lalam: That ability to return a specific piece of evidence as a potential support unit makes this incredibly useful for downstream applications because it gives us an immediate citation to check against.
Conclusion: Tom: So, wrapping up the discussion on "Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content, Evaluated Against Trained Verifiers," the main implication is that we can now audit RAG outputs at a granular level by measuring direct dependence on evidence rather than relying on aggregate scores.
Jane: It really reframes hallucination as a failure of dependence on evidence, which is a much more precise way to view the problem, and it provides an inexpensive guardrail for checking generated content in real-time.
Lu: The method’s strength lies in its ability to measure this relational property—how much a span depends on its evidence—rather than just looking at intrinsic text properties like fluency <ref:2607.04223#pg1>.
Meng: It remains competitive with chunk-level entailment verifiers, but it’s designed to be a scalable, self-contained scoring mechanism rather than requiring us to maintain and version an entirely separate verification model for every new application.
Lalam: Ultimately, GASP offers an explainable way to use the generator's own capabilities to flag unsupported claims, which is a significant step toward building more trustworthy generative AI systems across the board.
Mohamed Aly Bouke
Centre for Intelligent Cloud Computing, CoE for Advanced Cloud, Faculty of Information Science and Technology, Multimedia University
cs.CL, cs.AI, cs.LG
Submitted: 2026-07-05
Updated: 2026-10-02
Comments: 12 pages. Major revision and retitle of v1 (GASP, arXiv:2607.04223): recast as a controlled evaluation of a known with/without-context likelihood signal; results regenerated under a source-level split with identical inputs; adds an answer-level baseline, a cost analysis, and an annotator study. Code: https://github.com/drbouke/GASP
Code: https://github.com/drbouke/GASP
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 91/100
The gist: Retrieval-augmented generation (RAG) systems often suffer from hallucination, and this research introduces Grounding-Aware Sensitivity by Perturbation (GASP), a span-level detector that scores each
Key concepts
- Grounding Sensitivity
- This measures how much an answer sentence's probability changes when specific pieces of retrieved context are removed. It quantifies the dependence of a span on its supporting evidence, moving beyond simple response scoring to pinpoint unsupported claims.
- Random Nonlinear Iterated Function System (RNIFS)
- The decoding process under a given context is modeled as an RNIFS. This mathematical framework treats the set of grounded continuations as an attractor, allowing researchers to analyze how the system reacts when context units are perturbed or removed.
- Gap(y) and Drop(y)
- "Gap(y)" measures dependence on the entire context, while "Drop(y)" identifies the single most important context unit by measuring likelihood reduction when that specific chunk is deleted. These features combine to form a final hallucination score.
- Training-Free Detection
- GASP is designed to be used without any labeled data or training on verification models. It standardizes four grounding-sensitivity features and uses a simple thresholding mechanism (GASP-threshold) to flag suspicious spans, making it an inexpensive guardrail.
Terminology
Summary
Retrieval-augmented generation (RAG) systems often suffer from hallucination, and this research introduces Grounding-Aware Sensitivity by Perturbation (GASP), a span-level detector that scores each answer sentence by its grounding sensitivity—how strongly its likelihood depends on the retrieved evidence. This method is significant because it moves beyond simple response-level scoring to provide an explainable, training-free mechanism for identifying unsupported claims by measuring the dynamical system's reaction to context perturbation.
How it works
The core idea is to treat decoding under a context as a random nonlinear iterated function system (RNIFS)
whose attractor is the set of grounded continuations. Grounding sensitivity is defined as the distance between this invariant measure under full context and the measure when a specific context unit is removed, restricted to the answer span. This sensitivity is quantified using two complementary features:
-
The gap between the full-context and no-context log-likelihoods, denoted as
gap(y)
(Eq. 8), which measures dependence on the whole context. -
The maximum leave-one-out likelihood drop, denoted as
drop(y)
(Eq. 10), which localizes dependence to the single most important context unit by measuring the reduction in log-likelihood when that specific chunk is removed.
Detection and Scoring
The four grounding-sensitivity features are computed for each span: gap(y), jsd∅(y) (the mean JSD between full context and no-context predictions), drop(y), and jsdloo(y) (the maximum leave-one-out JSD). These features are aggregated to form a final hallucination score, which is defined as a decreasing function of them,
meaning higher scores indicate greater suspicion of hallucination. The method is training-free; the default detector, GASP-threshold, standardizes these four features and thresholds their negated sum, requiring no labeled data. A supervised variant, GASP-trained, uses a gradient-boosted decision tree classifier to map the feature vector to a hallucination score.
Interpretation and Explanation
The method provides an inherent explanation: for a flagged span, the unit attaining the maximum in Eq. (10) is returned as the candidate supporting passage.
For grounded spans, this identifies the single passage whose removal most reduced its likelihood,
serving as a candidate supporting unit rather than a verified citation. This mechanism is described through dynamical systems theory: removing evidence moves the context-conditioned invariant measure by an amount that reflects the span dependence on that removed unit.
Evaluation and Results
GASP was evaluated on three benchmarks (RAGTruth, TofuEval, RAGBench) using three instruction-tuned scorers from two model families (Qwen2.5-0.5B, Qwen2.5-1.5B, and SmolLM2-1.7B) under a leakage-clean protocol.
On RAGTruth, the span-level AUC reached approximately 0.67 to 0.73 across the scorers, significantly outperforming perplexity (AUC around 0.58 to 0.62) and length baselines (AUC near 0.55). The signal transferred effectively to TofuEval but was less pronounced on RAGBench short-answer question answering, suggesting GASP is best suited for outputs constructed from the retrieved context rather than answers recoverable from parametric knowledge.
Limitations and Use
The method's absolute AUC is moderate (0.64 to 0.75), indicating it is a signal that beats standard baselines rather than a finished product, and its performance on short-answer QA is bounded by perplexity. The primary use case is as an explainable guardrail and audit tool for RAG assistants,
highlighting low-sensitivity sentences for review or attaching the supporting passage to grounded sentences as an inline citation. The method measures a relational, counterfactual property—how much a span depends on its evidence—rather than intrinsic text properties like fluency.
Conclusion
Grounding sensitivity ranks hallucinated spans above grounded ones at both granularities and across all tested scorers, remaining competitive with a well-configured chunk-level entailment verifier while requiring no separate verifier model. The approach reframes hallucination as a failure of dependence on evidence rather than a property of the text, offering an explainable, inexpensive guardrail for retrieval-augmented systems.
The gist: Grounding sensitivity ranks hallucinated spans above grounded ones at both granularities and across all tested scorers, remaining competitive with a well-configured chunk-level entailment verifier while requiring no separate verifier model.
How it works
The core idea is to treat decoding under a context as a random nonlinear iterated function system (RNIFS)
whose attractor is the set of grounded continuations.
Improvements for AI systems
Based on the provided research article, Detecting Hallucinations in Retrieval-Augmented Generation through Grounding-Aware Sensitivity by Perturbation (GASP),
here are specific, actionable improvements for AI systems and what those improved systems can achieve:
The core improvement is shifting from coarse, response-level scoring to a fine-grained, evidence-based span-level detection mechanism.
Here are the specific improvements:
An implementable detector that scores each answer sentence based on its grounding sensitivity
(the change in likelihood when a specific context chunk is removed).
A training-free default detector utilizing a threshold on the standardized grounding-sensitivity features (derived from full context likelihoods, no-context likelihoods, and leave-one-out log-likelihood drops/JSD divergence).
An explainability mechanism that returns the specific retrieved context unit (chunk) that provided the strongest support for a grounded span. This provides an inline citation or supporting evidence for every flagged sentence, transforming the detector from a black-box score into an assistive guardrail.
A robust comparison framework that allows practitioners to compare grounding sensitivity directly against established baselines (perplexity, length, whole-context NLI entailment) without needing a separate, domain-specific verification model.
The improved AI system can do the following:
Flag specific sentences within a long RAG answer as hallucinated rather than flagging the entire response based on an aggregate score.
Provide users with immediate, actionable evidence (the supporting context chunk) to validate or refute a specific claim made by the LLM, dramatically improving trust in RAG outputs for high-stakes tasks like clinical or legal research.
Distinguish between two types of errors: baseless information
(details absent from the evidence) and evident conflict
(assertions contradicting specific retrieved passages), allowing downstream systems to prioritize which type of error is most dangerous.
Be deployed cheaply and efficiently as an inline guardrail, requiring only a small, probabilistic scorer rather than needing to retrain or maintain a separate, large verification model for every new application domain.
Abstract
Retrieval-augmented generation (RAG) assistants summarize records in clinical and legal work, where one unsupported sentence can mislead a reader. The contrast between an output's likelihood with and without its source is an established faithfulness score for whole summaries and answers, but it has not been measured as a detector of the individual unsupported sentence in multi-passage RAG answers, against trained verifiers, or for its cost. We implement it as a training-free detector that re-scores a fixed answer under the full context, no context, and each chunk removed, and returns the chunk whose removal lowers a sentence's likelihood most as a candidate supporting passage. We evaluate it on RAGTruth, TofuEval, and RAGBench with six scorers and against five verifiers, up to a large language model (LLM) judge, on identical inputs under a source-level split. Scoring per sentence ranks unsupported sentences better than the answer-level form of the same signal on all three benchmarks, by 0.033 to 0.071 in the area under the receiver operating characteristic curve (AUC). On RAGTruth the training-free score reaches an AUC of 0.717 to 0.745 across scorers and 0.773 with a classifier, above entailment and attribution baselines and level with per-chunk fact-checkers, at about one forty-seventh of the LLM judge's compute on a 1.5B scorer, while a full-context fact-checker and the judge are more accurate and are not improved by it. The signal is weakest on short-answer question answering, where the scorer can answer from memory.
Sources
- LettuceDetect: A Hallucination Detection Framework for RAG Applications
- RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering