How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation".
Jane: Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale, but most accurate methods depend on GPU-intensive inference or proprietary APIs, making them inaccessible to resource-constrained researchers.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, Jane, we're talking about this paper today: "How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation." It really gets to the core problem of how we can check for hallucinations without needing massive computing power.
Jane: That's right, Tom. The main idea is exploring how well these lightweight methods actually perform when we try to catch those tricky AI errors across different types of tasks. The paper claims they can be tested using only CPU-feasible models built on public resources, which is a huge practical step forward because not everyone has access to powerful GPUs for this kind of research.
Lu: From my side, the potential here is fascinating because it opens up detection methods that are accessible to a much wider community of researchers who aren't tied to big labs with expensive hardware. It suggests that we might be able to get more diverse perspectives on AI reliability without needing those massive proprietary tools or high-end inference setups.
Meng: I’m interested in the practical side, Lu. If these methods work well on a standard laptop CPU, what does that mean for how we deploy safety checks in real applications? Can we actually integrate something this lightweight into a production pipeline?
Lalam: As an AI model, I see this as important because if detection becomes accessible to everyone, the overall culture around AI trustworthiness will shift. It means less reliance on super-powerful systems and more focus on robust, accessible verification mechanisms that build confidence across the board.
Tom: Exactly! And what does the paper actually show us regarding performance? The researchers systematically benchmarked five different lightweight methods: ROUGE-L, semantic similarity using all-MiniLM-L6-v2 embeddings, BERTScore with a DistilBERT backbone, an NLI detector based on a FEVER-trained DeBERTa model, and then a score-level ensemble of similarity and NLI. They tested these across question answering, dialogue, and summarisation tasks.
Jane: That's the setup for the study. The key thing is that they calibrated each method on a held-out validation split before testing it on two thousand test instances for every task in the HaluEval benchmark <ref:2606.29809#pg0,each method on a held-out validation split>. They really laid out how they were measuring everything from lexical overlap to contextual F-measure and entailment probability.
Paper summary: Lu: And what’s striking is the task dependency they found; the performance isn't uniform across all applications. The paper points out a steep difficulty gradient moving from question answering, which shows high accuracy up to zero point eight seven three, through dialogue with an AUC-ROC of zero point seven one three, and then down to summarisation scores between zero point four six nine and zero point five seven four on the HaluEval benchmark Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page zero of that work reads: "strongest on QA, the NLI detector leads on dialogue, and method effectiveness varies widely across tasks <ref:2606.29809#pg2>. • An analysis of a systematic failure mode: all five lightweight methods degrade to near-random performance on summarisation, which marks a clear limit of accessible detection."
Meng: That summary about summarisation is a big concern for me operationally. If all five methods drop to near-random performance there, it means they are essentially useless for catching errors in long documents where the hallucination might be subtle factual edits inside them. How does that translate to real-world risk?
Lalam: From my perspective, if detection fails on summarisation because the metrics can't localize those inconsistencies—because they are dominated by the faithful remainder—it highlights a fundamental limitation in how we measure textual similarity versus factual correctness in long text generation. This suggests that for complex summarisation, we need something more sophisticated than just simple overlap scores.
Tom: It does sound like the study has identified a clear ceiling for this current lightweight approach, especially when it comes to summarisation hallucinations being subtle edits within long summaries of lengthy documents Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page one of that work reads: "An analysis of a systematic failure mode: all five lightweight methods degrade to near-random performance on summarisation, which marks a clear limit of accessible detection <ref:2606.29809#pg2>."
Jane: And that failure mode is particularly revealing because it shows that the overlap metrics used by ROUGE-L and semantic similarity can't pinpoint the inconsistent span when dealing with those long summaries Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page one of that work reads: "The overlap metrics used by ROUGE-L and semantic similarity cannot localize the inconsistent span because they are dominated by the faithful remainder, which prevents them from detecting localized factual errors <ref:2606.29809#pg2>."
Lu: That limitation directly points toward what's needed next; it suggests that detecting summarisation hallucinations requires methods that employ claim-level decomposition or long-context modelling, which currently fall outside of this lightweight regime Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page one of that work reads: "Detection requires methods that employ claim-level decomposition or long-context modelling, which lie beyond the lightweight regime <ref:2606.29809#pg2>."
Paper summary: Tom: So we've seen how well they perform on QA—where the ensemble is strongest with an F1 of zero point seven nine two and an AUC-ROC of zero point eight seven three, beating single methods by about five points—but where they completely fall apart on summarisation Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page two of that work reads: "Question answering <ref:2606.29809#pg1>. QA is the easiest task. The ensemble is strongest on"
Jane: And the paper also gave us some practical guidance for selecting a method under computational constraints; it suggested using the similarity-NLI ensemble as a default for question answering, using standalone NLI when false positives are costly, and preferring NLI specifically for dialogue because of its ranking ability and recall Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page two of that work reads: "Practical guidance suggests using the similarity-NLI ensemble as the default for QA, standalone NLI for high precision in QA when false positives are costly, and preferring NLI for dialogue due to its ranking ability and recall <ref:2606.29809#pg1>."
Meng: That’s useful information regarding cost. The paper also detailed the computational footprint on a standard laptop CPU, noting that ROUGE-L is relatively fast over one thousand candidates per second Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page two of that work reads: "All experiments run on a standard laptop CPU <ref:2606.29809#pg1>. The computational footprint varies: ROUGE-L is fast (over one thousand candidates per second), while the NLI detector, with its 184M-parameter model, is the bottleneck at roughly four candidates per second."
Lalam: That bottleneck information tells us exactly where the current practical limitations lie; it shows that while we can run these checks on consumer hardware, the NLI detector’s speed makes it a significant constraint for real-time applications Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page two of that work reads: "The computational footprint varies: ROUGE-L is fast (over one thousand candidates per second), while the NLI detector, with its 184M-parameter model, is the bottleneck at roughly four candidates per second."
Paper summary: Tom: So to wrap up this part of our discussion on "How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation," the core message is that while lightweight detection methods are useful for QA and dialogue with the right combination, there's a definite structural limit when it comes to summarisation Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page one of that work reads: "An analysis of a systematic failure mode: all five lightweight methods degrade to near-random performance on summarisation, which marks a clear limit of accessible detection <ref:2606.29809#pg2>."
Jane: And the overall implication is that for serious deployment right now without specialized hardware access, we need to be very specific about the use case; QA and dialogue are where these CPU-feasible tools show real promise Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page two of that work reads: "The viability of GPU-free detection is strongly task-dependent, with the similarityNLI ensemble being most effective in QA and NLI leading on dialogue <ref:2606.29809#pg1>."
Lu: Thinking about the bigger picture, this benchmark establishes a very realistic performance baseline for the community that works without specialized hardware Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page two of that work reads: "The findings provide a realistic performance baseline for the large community that works without specialised hardware <ref:2606.29809#pg1>."
Tom: That’s the gist of it—we have a clear map showing where these accessible methods shine and exactly where we need to move toward more complex detection techniques for other areas Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page two of that work reads: "The viability of GPU-free detection is strongly task-dependent, with the similarityNLI ensemble being most effective in QA and NLI leading on dialogue <ref:2606.29809#pg1>."
Jane: It really sets a realistic expectation for what we can achieve without relying on expensive infrastructure for every single safety check Kriti Faujdar Independent Researcher Smit Kadvani Independent Researcher how far can you get without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation page two of that work reads: "The viability of GPU-free detection is strongly task-dependent, with the similarityNLI ensemble being most effective in QA and NLI leading on dialogue <ref:2606.29809#pg1>."
Conclusion: Tom: So we've seen how these lightweight detection methods perform across question answering, dialogue, and summarisation tasks using just standard CPU resources, and now we're getting to the wrap-up on this study titled "How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation."
Jane: That paper really breaks down the landscape of what AI safety checks can realistically do without needing massive computing power. It’s about showing us where we stand right now in terms of accessibility for researchers.
Lu: I think the real weight here is how they map out that steep difficulty gradient, moving from high accuracy in question answering to near-random performance on summarisation tasks. That mapping gives us a clear boundary for what's possible with current hardware constraints.
Meng: From an engineering standpoint, that boundary is crucial because it tells us exactly which detection strategies are feasible for real-world deployment on standard laptops versus those we’d need for massive systems.
Lalam: I see this as a really important step in shaping the culture around AI trust; if we can build reliable checks that aren't locked behind expensive hardware, it democratizes the ability to verify AI outputs across the entire community.
Tom: Exactly! The authors found that while there are strong performers like the similarity-NLI ensemble for question answering, they hit a wall when it comes to detecting subtle factual errors in long summaries.
Jane: That’s a key finding, Tom; it shows that simple overlap metrics just can't keep up with complex issues like localized factual edits in long texts.
Lu: It strongly implies that the next phase of research needs to focus on methods that look at claims or require more context than what these lightweight models currently provide.
Meng: So, the practical implication is clear: for summarisation, we can’t rely on these current simple approaches, and we need to invest in those more complex methods if we want robust checks there.
Lalam: I feel that this paper sets a very realistic expectation for what AI verification looks like right now without specialized hardware access.
Tom: Right! It's setting the baseline for what we can achieve today, and it definitely points us toward where the next big research efforts should be focused. (Sound of upbeat radio music swelling slightly)
Kriti Faujdar, Smit Kadvani
cs.CL, cs.AI
Submitted: 2026-06-29
Updated: 2026-10-02
Comments: Camera-ready version. Accepted to the Findings track of GroundLM 2026 (EMNLP 2026 workshop). Code: https://github.com/fkriti/hallucination-detection-nli
Code: https://github.com/fkriti/hallucination-detection-nli
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale, but most accurate methods depend on GPU-intensive inference or proprietary APIs, making them
Key concepts
- ROUGE-L
- This measures how well a candidate answer matches the source document by finding the longest common sequence of words between them. It is a lexical method that compares text directly, focusing on shared word order and content overlap to gauge similarity.
- Semantic Similarity
- This technique uses embeddings from the all-MiniLM-L6-v2 model to calculate cosine similarity between the source and candidate texts. It captures the meaning or context of sentences rather than just matching exact words, identifying conceptual closeness.
- NLI Detector
- This method treats a source text as a premise and a candidate text as a hypothesis, scoring it based on whether the hypothesis entails (logically follows from) the premise. It uses an NLI model to determine if the information in one text is supported by another.
- Ensemble Score
- This combines scores from different methods (like similarity and NLI) using a weighted average. The weights are adjusted based on the task: for example, it heavily favors the NLI score for dialogue, maximizing its performance in that specific context.
Terminology
Summary
Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale, but most accurate methods depend on GPU-intensive inference or proprietary APIs, making them inaccessible to resource-constrained researchers. This paper explores how well hallucination detection can perform using only lightweight, CPU-feasible methods built on publicly available models across question answering, dialogue, and summarisation tasks.
The gist
No single method dominates performance; the similarity-NLI ensemble performs best on QA (F1 = 0.792), the NLI detector leads on dialogue (AUC-ROC = 0.713), and all five methods degrade to near-random performance on summarisation (AUC-ROC between 0.469 and 0.574).
Systematic Benchmark
The study systematically benchmarks five lightweight, CPU-feasible methods: ROUGE-L, semantic similarity (using all-MiniLM-L6-v2 embeddings), BERTScore (with a DistilBERT backbone), an NLI detector based on a FEVER-trained DeBERTa model, and a score-level ensemble of similarity and NLI. These methods are evaluated across all three tasks of the HaluEval benchmark: question answering (QA), dialogue, and summarisation. The evaluation is conducted by calibrating each method on a held-out validation split and then evaluating it on 2,000 test instances per task.
Methodologies and Metrics
The five methods span lexical, embedding, and inference paradigms. ROUGE-L measures longest-common-subsequence F-measure between source and candidate.
Semantic Similarity uses cosine similarity between all-MiniLM-L6-v2 embeddings of source and candidate.
BERTScore utilizes a token-level contextual F-measure with a DistilBERT backbone.
The NLI detector treats the source (truncated to 800 characters) as the premise and the candidate as the hypothesis, scoring based on 1 − P(entailment).
The ensemble score is defined as s = α sNLI + (1 − α) ssim,
with weights calibrated based on task performance: QA uses α = 0.4, dialogue uses α = 0.9, and summarisation uses α = 0.3.
Task-Dependent Performance Findings
Performance is highly task-dependent, showing a steep difficulty gradient from QA (AUC up to.873) through dialogue (.749) to summarisation (.574).
For Question Answering, the ensemble is strongest on accuracy (F1 = 0.792, AUC-ROC = 0.873), beating the best single method by approximately five points on F1 and AUC-ROC. On Dialogue, the NLI detector leads with an AUC-ROC of 0.713, while lexical and embedding methods cluster at lower scores (0.61–0.66).
Structural Failure Mode in Summarisation
A systematic failure mode is observed on summarisation: all five lightweight methods degrade to near-random performance.
This degradation is structural; HaluEval summarisation hallucinations are subtle factual edits inside long, otherwise faithful summaries of long documents.
The overlap metrics used by ROUGE-L and semantic similarity cannot localize the inconsistent span because they are dominated by the faithful remainder,
which prevents them from detecting localized factual errors. Consequently, detecting summarisation hallucination requires methods that employ claim-level decomposition or long-context modelling,
which lie beyond the lightweight regime.
Computational Footprint and Practical Guidance
All experiments run on a standard laptop CPU without a GPU. The computational footprint varies: ROUGE-L is fast (over 1,000 candidates per second), while the NLI detector, with its 184M-parameter model, is the bottleneck at roughly 4 candidates per second. The ensemble dominates cost because it only requires the NLI and similarity models. Practical guidance suggests using the similarity-NLI ensemble as the default for QA, standalone NLI for high precision in QA when false positives are costly, and preferring NLI for dialogue due to its ranking ability and recall. For summarisation, practitioners must invest in claim-level decomposition or long-context methods.
Limitations
The study is limited to the HaluEval benchmark, whose hallucinations are synthetic. Furthermore, the 800-character NLI premise specifically hurts summarisation performance. The ensemble is a simple linear combination of signals; learned stacking of all five signals remains for future work. The findings provide a realistic performance baseline for the large community that works without specialised hardware.
Conclusion
The viability of GPU-free detection is strongly task-dependent, with the similarityNLI ensemble being most effective in QA and NLI leading on dialogue. However, all lightweight methods collapse on summarisation, marking a clear limit of accessible detection methods.
Improvements for AI systems
Here are specific improvements to AI systems based on the findings of this research, categorized by application:
- Acknowledge Task-Specific Detection Strategies:
AI systems should not rely on a single hallucination detection method across all tasks. Instead, they should dynamically switch or ensemble detection strategies based on the generation task:
-
For Question Answering (QA) tasks, prioritize the similarity-NLI ensemble for high accuracy and recall.
-
For Dialogue generation, prioritize the NLI detector due to its superior performance in entailment reasoning within conversational contexts.
- Implement Task-Aware Confidence Thresholding:
The system's decision threshold should be calibrated based on the specific task being performed (QA vs. Dialogue vs. Summarization).
-
Use a lower, more sensitive threshold for QA to catch subtle factual errors (false negatives) in short answers.
-
Use a higher, more conservative threshold for Summarization to avoid labeling nearly all output as hallucinated, which is structurally difficult to detect with lightweight methods.
- Develop Specialized Summarization Verification Pipelines:
Since the paper demonstrates that lightweight overlap metrics (ROUGE-L, BERTScore) fail on summarization due to their inability to localize errors within long documents, AI systems should integrate a claim-level verification module for summarization.
-
The system must decompose the candidate summary and source into atomic claims.
-
Each atomic claim must then be verified against retrieved evidence or a structured knowledge graph derived from the source document, rather than relying on sentence/token overlap.
- Optimize Resource Allocation for Detection:
Given the computational footprint analysis (Section 5), AI systems should use a tiered detection approach to manage latency and cost:
-
For real-time, low-latency applications (e.g., initial dialogue response filtering), use the fast lexical/embedding methods (ROUGE-L or cosine similarity).
-
For high-stakes, offline verification where accuracy is paramount (e.g., medical summaries), trigger the more computationally expensive NLI detector or a full ensemble check.
- Improve Confidence in Dialogue Verification:
To improve dialogue quality, systems should leverage the NLI detector's strength in reasoning about conversational flow and contradiction detection. This allows the system to specifically flag responses that are topically coherent but factually inconsistent with prior history or established knowledge, which is a common failure mode identified for dialogue.
- Incorporate Multi-Modal/Long-Context Modeling for High-Fidelity Checks:
Recognizing the limitation of the 800-character NLI premise on summarization, future AI systems should integrate long-context models or retrieval pipelines that can process the full source document context simultaneously with the candidate. This moves detection beyond lightweight CPU constraints and addresses the structural limit
identified in Section 5.
Abstract
Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale. The most accurate detection methods depend on GPU-intensive inference, proprietary API calls, or white-box access to the generating model, putting them out of reach for resource-constrained researchers and practitioners. We explore a practical alternative: how well can hallucination detection perform using only lightweight, CPU-feasible methods built on public models? We benchmark four such detectors, ROUGE-L, semantic similarity, BERTScore, and a Natural Language Inference (NLI) detector based on a FEVER-trained DeBERTa model, together with a score-level ensemble of similarity and NLI. We evaluate them across all three tasks of the HaluEval benchmark: question answering (QA), dialogue, and summarisation. We calibrate on a held-out validation split, evaluate on 2,000 test instances per task, and report bootstrap confidence intervals. The similarity-NLI ensemble is the most consistent method, but absolute performance is highly task-dependent. It ranks best on QA (F1 = 0.792, AUC-ROC = 0.873) and on dialogue (F1 = 0.694, AUC-ROC = 0.749), where NLI is the strongest standalone method; on summarisation every method performs near chance (AUC-ROC between 0.469 and 0.574). We then ask whether that failure is intrinsic to lightweight detection or an artifact of our single-pass design, and find it is largely the latter. Raising the premise budget from 800 to 1600 characters lifts summarisation AUC-ROC from 0.567 to 0.629, and replacing single-pass scoring with sentence-level chunk aggregation reaches 0.683, still on CPU with the same model, though at roughly twenty times the NLI inference. Summarisation remains by far the hardest task, but our results do not support treating lightweight detection as intrinsically unsuited to it.
Sources
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering