"You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused
cs.AI, cs.CR
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/prevalentWare/opencod
Terminology
Sources
- Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models
- Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
- The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models
- Do System Prompts Leave Behavioral Fingerprints? A Large-Scale Empirical Study of Clone Detection via Output Similarity
- It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
- Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security
- Gotta Catch them all: the modes of Sycophancy
- Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse
- Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
- ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
- SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
- Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
- Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH
- LLMs Corrupt Your Documents When You Delegate
- AgentCanary: A Security Evaluation Framework for Autonomous AI Agents in Real Executable Environments
- Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
- OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills
- ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents
- Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Self-Improving Personal Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection