AEScorer: An Agentic Evidence-Grounded Framework for Graded Factuality Verification
cs.CL
Submitted: 2026-01-07
Updated: 2026-08-27
Terminology
Sources
- FELM: Benchmarking Factuality Evaluation of Large Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- LoRA: Low-Rank Adaptation of Large Language Models
- FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios
- Learning to Verify Summary Facts with Fine-Grained LLM Feedback
- A Revisit of Fake News Dataset with Augmented Fact-checking by ChatGPT
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- A Comprehensive Survey on Long Context Language Modeling
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
- A Survey of Context Engineering for Large Language Models
- HalluciNot: Hallucination Detection Through Context and Common Knowledge Verification
- Towards Reliable Misinformation Mitigation: Generalization, Uncertainty, and GPT-4
- Atomic Inference for NLI with Generated Facts as Atoms
- LLaMA: Open and Efficient Foundation Language Models
- Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Qwen3 Technical Report
- A Survey of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering