AgentPack: A Dataset of Code Changes, Co-Authored by Agents and Humans
cs.SE, cs.CL
Submitted: 2025-09-26
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Fine-tuning large language models for code editing has typically relied on mining commits and pull requests.
Terminology
Abstract
Fine-tuning large language models for code editing has typically relied on mining commits and pull requests. The working hypothesis has been that commit messages describe human intent in natural language, and patches to code describe the changes that implement that intent. However, much of the previously collected data is noisy: commit messages are terse, human-written commits commingle several unrelated edits, and many commits come from simple, rule-based bots. The recent adoption of software engineering agents changes this landscape. Code changes co-authored by humans and agents are often accompanied by substantially more explicit natural-language descriptions of intent and rationale. Moreover, when these changes land in public repositories, they are implicitly filtered by humans: maintainers discard low-quality commits to their projects. We present AgentPack, a corpus of 1.8M code edits co-authored by Claude Code, OpenAI Codex, and Cursor Agent across public GitHub projects up to early October 2025. We describe the identification and curation pipeline, quantify adoption trends of these agents, and analyze the structural properties of the edits. Finally, we show that models fine-tuned on AgentPack can outperform models trained on prior human-only commit corpora, highlighting the potential of using public data from software engineering agents to train future code-editing models.
Sources
- Multi-Turn Code Generation Through Single-Step Rewards
- Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP
- Program Synthesis with Large Language Models
- Agnostics: Learning to Code in Any Programming Language via Reinforcement with a Universal Learning Environment
- Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs
- Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions
- Evaluating Large Language Models Trained on Code
- OctoPack: Instruction Tuning Code Large Language Models
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
- SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
- Magicoder: Empowering Code Generation with OSS-Instruct
- Qwen3 Technical Report
- ACECODER: Acing Coder RL via Automated Test-Case Synthesis
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties