Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 22 pages, 7 figures. Extended version of a paper accepted at EvalLLM 2025 (CORIA-TALN 2025)
Code: https://github.com/GiovanniGatti/truthbench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Evaluating the factual correctness of large language models (LLMs) is vital for many applications.
Terminology
Abstract
Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factuality-based metrics, their sensitivity and reliability remain underexplored. This paper introduces a meta-evaluation framework that systematically tests these metrics using controlled corruptions of gold standard answers. Our method generates ranked outputs with known degrees of degradation to probe how metrics capture nuanced changes in truthfulness. Our experiments reveal that pipeline-based methods, such as the RAGAS's factual correctness metric, better track degradation than LLM-as-judge approaches. We also propose a new variant of the factual correctness metric that provides a competitive and cost-efficient.
Sources
- FELM: Benchmarking Factuality Evaluation of Large Language Models
- SummEval: Re-evaluating Summarization Evaluation
- A Survey on LLM-as-a-Judge
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- JudgeBench: A Benchmark for Evaluating LLM-based Judges
- Evaluating Large Language Models at Evaluating Instruction Following
- BERTScore: Evaluating Text Generation with BERT
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering