SLMFix: Leveraging Small Language Models for Domain Specific Language Error Fixing with Reinforcement Learning
cs.SE, cs.AI, cs.PL
Submitted: 2025-11-24
Updated: 2026-09-25
Code: https://github.com/ansible/ansible-content-parser
Terminology
Sources
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- MultiCoder: Multi-Programming-Lingual Pre-Training for Low-Resource Code Completion
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- Qwen2.5-Coder Technical Report
- Evaluating Large Language Models Trained on Code
- Towards Formal Verification of LLM-Generated Code from Natural Language Prompts
- Coarse-Tuning Models of Code with Reinforcement Learning Feedback
- RLSF: Fine-tuning LLMs via Symbolic Feedback
- Benchmarking Educational Program Repair
- How Small is Enough? Empirical Evidence of Quantized Small Language Models for Automated Program Repair
- Competition-Level Code Generation with AlphaCode
- StarCoder 2 and The Stack v2: The Next Generation
- Granite Code Models: A Family of Open Foundation Models for Code Intelligence
- GPT-4 Technical Report
- Competitive Programming with Large Reasoning Models
- Proximal Policy Optimization Algorithms
- HybridFlow: A Flexible and Efficient RLHF Framework
- SLM-SQL: An Exploration of Small Language Models for Text-to-SQL
- From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
- Qwen2.5 Technical Report
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties