Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs
cs.CL, cs.LG
Submitted: 2026-09-03
Updated: 2026-09-03
Comments: 21 pages, 7 figures. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) through executable feedback.
Terminology
Abstract
Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) through executable feedback. The feedback for coding problems mainly comes from specific test cases, where high-quality test cases are often scarce since they should be both sound and discriminative. We thus turn to study the auto-generation of test cases using the learned model. We find this is naturally an adversarial RL problem: the model is expected to generate effective test cases as counterexamples, depending on the solver's current failure modes. We propose Test Cases Scaling (TCS), a two-stage RL framework for effective test generation. Both stages train a test generator from a rolling policy-aligned buffer: Stage 1 generates tests consistent with the reference solution, and Stage 2 restricts the buffer to current failure modes and learns counterexample tests. Across TACO and LiveCodeBench, TCS improves both pass@1 and inference-time answer selection according to generated tests. We find the learned test generator also enables effective selection among other LLM outputs.
Sources
- InternLM2 Technical Report
- A Survey on Code Generation with LLM-based Agents
- Reward Shaping to Mitigate Reward Hacking in RLHF
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- Skywork Open Reasoner 1 Technical Report
- Qwen2.5-Coder Technical Report
- Large Language Models as Test Case Generators: Performance Evaluation and Enhancement
- TACO: Topics in Algorithmic COde generation dataset
- Learning to Solve and Verify: A Self-Play Framework for Code and Test Generation
- Inference-Time Scaling for Generalist Reward Modeling
- Code Llama: Open Foundation Models for Code
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering