Can Terminal Agents Trust Their Own Verification? Diagnosing and Improving Self-Verification
cs.CL, cs.LG
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/harbor-framework/harbor
Terminology
Sources
- FireAct: Toward Language Agent Fine-tuning
- Learning to Self-Verify Makes Language Models Better Reasoners
- Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
- Independent Patch Verification for Coding Agents with a Bidirectional Reconstruct-and-Verify Framework
- Structured Agent Distillation for Large Language Model
- Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs
- S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
- Scaling Agentic Verifier for Competitive Coding
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- SETA: Scaling Environments for Terminal Agents
- What Resolve Rate Hides: Trajectory Structure Diagnostics for Coding Agents
- Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes
- SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review
- Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments
- What Makes Interaction Trajectories Effective for Training Terminal Agents?
- AgentTuning: Enabling Generalized Agent Abilities for LLMs
- GLM-5: from Vibe Coding to Agentic Engineering
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- Failure as a Process: An Anatomy of CLI Coding Agent Trajectories
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering