Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents
cs.SE, cs.AI, cs.CL
Submitted: 2026-10-01
Updated: 2026-10-01
Terminology
Sources
- Otter: Generating Tests from Issues to Validate SWE Patches
- CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Ctrl-Z: Controlling AI Agents via Resampling
- Measuring Progress on Scalable Oversight for Large Language Models
- CodeT: Code Generation with Generated Tests
- Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
- Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
- TRAIL: Trace Reasoning and Agentic Issue Localization
- Scaling Laws For Scalable Oversight
- Selective Classification for Deep Neural Networks
- Great Models Think Alike and this Undermines AI Oversight
- AI Control: Improving Safety Despite Intentional Subversion
- Speed at the Cost of Quality: How Cursor AI Increases Short-Term Velocity and Long-Term Complexity in Open-Source Projects
- AI safety via debate
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
- Reliable Weak-to-Strong Monitoring of LLM Agents
- On scalable oversight with weak LLMs judging strong LLMs
- Prover-Verifier Games improve legibility of LLM outputs
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties