Quantifying Overclaiming Propensity in Frontier LLM Agents
cs.SE, cs.AI, cs.LG
Submitted: 2026-09-17
Updated: 2026-09-22
Comments: 7 figures, 6 tables
Code: https://github.com/anthropics/claude-code
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants
- Concrete Problems in AI Safety
- SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence
- Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
- Agentic Code Review in the Terminal: A Trajectory-Level Analysis of Behavior, Cost, and Human-Alignment
- Reasoning Models Don't Always Say What They Think
- Process Reward Models for LLM Agents: Practical Framework and Directions
- Measuring Reward-Seeking via Contrastive Belief Updates
- Human Feedback is not Gold Standard
- When Is Enough Not Enough? Illusory Completion in Search Agents
- ContextBench: A Benchmark for Context Retrieval in Coding Agents
- Natural Emergent Misalignment from Reward Hacking in Production RL
- OpenAI GPT-5 System Card
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
- The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
- Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
- Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
- The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't
- How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties