Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
cs.CL
Submitted: 2025-10-09
Updated: 2026-05-07
Comments: ACL 2026 Findings
Journal ref: Findings of the Association for Computational Linguistics: ACL 2026, 34931-34966. 2026
DOI: 10.18653/v1/2026.findings-acl.1744
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
- FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios
- The Llama 3 Herd of Models
- VeriFact: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference Facts
- Measuring short-form factuality in large language models
- Qwen2.5 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering