Mitigating Hallucination in Large Language Models: A Capability-Oriented Survey on RAG, Reasoning, and Agentic Systems
cs.CL, cs.AI
Submitted: 2025-10-28
Updated: 2026-09-27
Terminology
Sources
- Medical Hallucinations in Foundation Models and Their Impact on Healthcare
- GPT-4 Technical Report
- DeepSeek-V3 Technical Report
- Mistral 7B
- A Survey on Large Language Model Hallucination via a Creativity Perspective
- Gemini: A Family of Highly Capable Multimodal Models
- L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit
- InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
- MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
- Agent-UniRAG: A Trainable Open-Source LLM Agent Framework for Unified Retrieval-Augmented Generation Systems
- MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
- HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation
- Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision
- AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
- BatGPT: A Bidirectional Autoregessive Talker from Generative Pre-trained Transformer
- A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
- LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain
- Code Llama: Open Foundation Models for Code
- The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering