When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
cs.AI, cs.CL, cs.LG
Submitted: 2026-05-28
Updated: 2026-08-31
Terminology
Sources
- MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents
- Training Verifiers to Solve Math Word Problems
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure
- CL-bench: A Benchmark for Context Learning
- A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback
- CUB: Benchmarking Context Utilisation Techniques for Language Models
- Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
- Breaking Contextual Inertia: Reinforcement Learning with Single-Turn Anchors for Stable Multi-Turn Interaction
- Do LLMs Benefit From Their Own Words?
- BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents
- Qwen2.5 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Meta-Harness: End-to-End Optimization of Model Harnesses
- CaRT: Teaching LLM Agents to Know When They Know Enough
- OpenAI GPT-5 System Card
- Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
- Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- Hallucinations Undermine Trust; Metacognition is a Way Forward
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection