Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict
cs.AI, cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/KaiserWhoLearns/EpistemicHumilityLLMAgents
Terminology
Sources
- Measuring Progress on Scalable Oversight for Large Language Models
- AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security
- TRAIL: Trace Reasoning and Agentic Issue Localization
- How to Interpret Agent Behavior
- Monitoring Monitorability
- Training LLMs for Honesty via Confessions
- Language Models (Mostly) Know What They Know
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- A Survey on the Honesty of Large Language Models
- Teaching Models to Express Their Uncertainty in Words
- Towards Understanding Sycophancy in Language Models
- OpenAI GPT-5 System Card
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
- ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM
- Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict
- Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
- Measuring Epistemic Humility in Multimodal Large Language Models
- Retrieval-Augmented Generation with Conflicting Evidence
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- Know Your Limits: A Survey of Abstention in Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection