Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
cs.AI, cs.CL
Submitted: 2026-08-27
Updated: 2026-09-25
Code: https://github.com/Pranav-1100/confidence-calibration-evaluation
Terminology
Sources
- Distinguishing the Knowable from the Unknowable with Language Models
- QLoRA: Efficient Finetuning of Quantized LLMs
- Language Models (Mostly) Know What They Know
- Why Language Models Hallucinate
- ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- AgentAbstain: Do LLM Agents Know When Not to Act?
- Agentic Abstention: Do Agents Know When to Stop Instead of Act?
- Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions
- Qwen2.5 Technical Report
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- LLM Agents Already Know When to Call Tools -- Even Without Reasoning
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
- Cherry-pick Override: LLM Judges Under-use the Non-Directional Verdicts Their Contract Authorizes
- The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents
- Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection