Hallucination in Multimodal Foundation Models: A Survey on Causes, Corrections, and Evaluations
cs.AI
Submitted: 2024-10-20
Updated: 2026-09-13
Code: https://github.com/lyq312318224/MLLMsAugmented
Project page: https://fuxiaoliu.github.io/LRV/https://github.com/Yuqifan1117/HalluciDoctor
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- OphGLM: Training an Ophthalmology Large Language-and-Vision Assistant based on Instructions and Dialogue
- ADriver-I: A General World Model for Autonomous Driving
- Improved Baselines with Visual Instruction Tuning
- MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- Hallucination of Multimodal Large Language Models: A Survey
- A Survey on Hallucination in Large Vision-Language Models
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
- OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
- Scaling Instruction-Finetuned Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- MLLMs-Augmented Visual-Language Representation Learning
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- CIEM: Contrastive Instruction Evaluation Method for Better Instruction Tuning
- FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing
- Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection