Steer, Don't Solve: Training Small Critic Models for Large Code Agents
cs.SE, cs.AI, cs.LG
Submitted: 2026-06-20
Updated: 2026-09-01
Code: https://github.com/shubhamrgandhi/critic-training
License: http://creativecommons.org/licenses/by/4.0/
The gist: Coding tasks are typically complicated and require multiple capabilities, ranging from high-level planning to low-level implementation.
Terminology
Abstract
Coding tasks are typically complicated and require multiple capabilities, ranging from high-level planning to low-level implementation. While coding agents are optimized for the joint capabilities, individual capabilities such as high-level planning may have different optima and remain a major bottleneck. To address this challenge, we train a separate critic model that is specialized in high-level planning to steer the coding agent in inference. We construct SFT and DPO data to train the critic model to identify errors made by the coding agent and provide correct and clear high-level guidance without generating concrete actions. Experiments show that our fine-tuned 4B and 8B critic models significantly improve the performance of 6 larger coding agents (e.g., improving the resolved rates of GLM-4.7-Flash-30B-A3B and GPT-OSS-120B by 16.0% and 14.4% on SWE-Bench Verified). The critic model also reduces the total inference costs for some coding agents by solving tasks in fewer steps (e.g., reducing the per-example inference cost for GPT-OSS-20B from 0.07 to 0.03). Code: https://github.com/shubhamrgandhi/critic-training
Sources
- MASAI: Modular Architecture for Software-engineering AI Agents
- How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
- When Agents go Astray: Course-Correcting SWE Agents with PRMs
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- CWM: An Open-Weights LLM for Research on Code Generation with World Models
- A Rubric-Supervised Critic from Sparse Real-World Outcomes
- SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
- An Empirical Study on Failures in Automated Issue Solving
- Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
- LLM Critics Help Catch LLM Bugs
- ProRefine: Inference-Time Prompt Refinement with Textual Feedback
- Large Language Model Critics for Execution-Free Evaluation of Code Changes
- Qwen3 Technical Report
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties