HALT: Hallucination Assessment via Log-probs as Time series
cs.CL, cs.AI
Submitted: 2026-02-02
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Hallucinations remain a major obstacle for large language models (LLMs), especially in safety-critical domains.
Terminology
Abstract
Hallucinations remain a major obstacle for large language models (LLMs), especially in safety-critical domains. We present HALT (Hallucination Assessment via Log-probs as Time series), a lightweight hallucination detector that leverages only the top-20 token log-probabilities from LLM generations as a time series. HALT uses a gated recurrent unit model combined with entropy-based features to learn model calibration bias, providing an extremely efficient alternative to large encoders. Unlike white-box approaches, HALT does not require access to hidden states or attention maps, relying only on output log-probabilities. Unlike black-box approaches, it operates on log-probs rather than surface-form text, which enables stronger domain generalization and compatibility with proprietary LLMs without requiring access to internal weights. To benchmark performance, we introduce HUB (Hallucination detection Unified Benchmark), which consolidates prior datasets into ten capabilities covering both reasoning tasks (Algorithmic, Commonsense, Mathematical, Symbolic, Code Generation) and general purpose skills (Chat, Data-to-Text, Question Answering, Summarization, World Knowledge). While being 30x smaller, HALT outperforms Lettuce, a fine-tuned modernBERT-base encoder, achieving a 60x speedup gain on HUB. HALT and HUB together establish an effective framework for hallucination detection across diverse LLM capabilities.
Sources
- INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Calibration of Pre-trained Transformers
- RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
- The Llama 3 Herd of Models
- On Calibration of Modern Neural Networks
- OpenAssistant Conversations -- Democratizing Large Language Model Alignment
- CriticEval: Evaluating Large Language Model as Critic
- Contrastive Learning Reduces Hallucination in Conversations
- Revisiting the Calibration of Modern Neural Networks
- Fine-grained Hallucination Detection and Editing for Language Models
- DecoPrompt : Decoding Prompts Reduces Hallucinations when Large Language Models Meet False Premises
- Qwen3 Technical Report
- RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
- Hallucination Detection in Large Language Models with Metamorphic Relations
- Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
- LettuceDetect: A Hallucination Detection Framework for RAG Applications
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering