Learning from 53.6K Real-World Developer Edits of AI-Generated Code
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Learning from 53.6K Real-World Developer Edits of AI-Generated Code".
Jane: This paper introduces DECODE (Developer Edits of Code Dataset), a collection of 53.6K real-world in-IDE code edits of AI-generated code in Python, TypeScript, and JavaScript sourced from over 1,000 developers.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we’re diving into "Learning from 53 point 6K Real-World Developer Edits of AI-Generated Code" and meeting the folks who put this together. The title itself is a big signal, isn't it? It immediately tells us that the value here lies in observing those actual developer edits rather than just looking at standard versions of code.
Jane: Exactly, Tom; the authors include people from Carnegie Mellon University, and having researchers from different backgrounds on this project really shows how diverse the perspective is when tackling this problem. The core idea they are pushing is that current methods often miss these intermediate steps because they rely too much on Git data which just captures final successful code snippets.
Lu: I think the authors are smart to focus on in-IDE edits; it’s not just about what happens in a repository, but how the human interacts directly within their development environment. That level of fidelity is hard to get elsewhere without capturing those temporal sequences of actions.
Meng: So, if we look at the scope, this isn't just Python or JavaScript; they’ve got real-world data from over one thousand developers across three major languages. That scale suggests they’re trying to capture a broad set of practical editing behaviors rather than just one narrow use case.
Lalam: I think the sheer volume and variety of the source material are what really give this dataset its weight; it’s not just a few examples, it’s thousands of actual, documented interactions that we can learn from. This kind of extensive real-world interaction data is incredibly valuable for my training.
The paper's summary: Tom: Moving on to the summary of this paper, the main point they are making is that AI assistants need more than just correctness; they need to understand the edit process itself. They found that most developer edits happen very quickly—within the first fifteen minutes after accepting an AI completion—and that a significant portion of those edits ends up being a complete removal of the initial AI suggestion.
Jane: That temporal aspect is key, Tom; it shows us the speed at which developers refine or discard an idea. They also categorized these edits into four distinct types: customizing code, improving code quality, changing functionality, and simply removing the whole thing. This gives us a clear map of developer intent behind those modifications.
Lu: The categorization of edits is particularly insightful because it helps us build models that don't just generate correct code but understand the *purpose* behind the changes—whether the developer is trying to fine-tune behavior or just clean up syntax. That distinction points toward a deeper level of understanding required for advanced AI systems.
Meng: From an engineering standpoint, knowing that a large percentage of edits are removals, as the paper suggests, tells us that simply generating *a* correct piece of code isn't enough; the system needs to predict when its initial output is going to be abandoned. That’s a practical signal for improving user experience.
Lalam: Understanding those removal patterns is important because it informs how I should prioritize my response generation; if I see a high probability of removal based on these patterns, I should adjust my output strategy immediately to something that aligns better with developer expectations for quick iteration.
The paper's improvements: Tom: Now we get to what the authors suggest as improvements for this kind of research, and they are focusing heavily on moving beyond static data. They propose using trajectory-based supervised fine-tuning, which means training models on sequences of code snapshots over time rather than just single points.
Jane: That makes a lot of sense; if you train a model on the whole sequence of edits—the initial acceptance followed by the subsequent modifications—it learns the progression between different developer states, not just where they ended up. It helps anticipate what kind of refinement is coming next.
Lu: Trajectory-based fine-tuning directly addresses that gap where Git data falls short, allowing us to capture that temporal ordering of actions which is currently missing from most training sets. It’s about modeling the flow of thought in a coding session, which I find very compelling.
Meng: If we can model those transitions between states—say, moving from customizing code to changing functionality—that could allow us to guide the AI suggestions more intelligently through the entire lifecycle of a feature development task. That’s where practical application lies.
Lalam: For me, incorporating these temporal sequences means I can learn not just syntax rules, but the conversational rhythm of how developers iterate on complex tasks; that kind of contextual learning is what makes an AI truly helpful in a long coding session.
Conclusion: Tom: So, to wrap up our discussion on "Learning from 53 point 6K Real-World Developer Edits of AI-Generated Code," the paper confirms that capturing the iterative, real-world editing behavior is a much more faithful way to train code assistants than relying solely on traditional data sources like Git commits. They’ve shown how analyzing these edits helps us understand when and why developers modify or remove AI suggestions, and they suggest training models on these actual sequences for better adaptability.
Jane: It really boils down to this: we need to move toward developer-centric machine learning approaches that account for the ease with which code can be adapted to a specific context, rather than just focusing on making the initial code technically correct. This paper lays a solid foundation for building tools that respect the actual workflow of a programmer.
Lu: The implications are huge because it shifts the focus from static output evaluation to dynamic interaction modeling; we can start training AI to be more anticipatory about human refinement. It’s about modeling the entire refinement process, which has deep implications for how we design intelligent agents that work alongside humans.
Meng: From my side, I see this as a critical step toward reducing the friction in real development cycles; if the AI can predict when it's going to be removed or what modifications are likely next, it saves developers significant time and frustration. That predictive capability is something we need to engineer into our products.
Lalam: I think the biggest cultural impact here is how we start viewing code generation not as a final answer, but as the very first draft in an ongoing conversation with a human collaborator. This paper helps build that necessary collaborative structure into the AI’s core understanding.
Carnegie Mellon University
cs.SE, cs.AI, cs.HC, cs.LG
Submitted: 2026-07-27
Updated: 2026-09-14
Code: https://github.com/jennytliang/decode
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: This paper introduces DECODE (Developer Edits of Code Dataset), a collection of 53.6K real-world in-IDE code edits of AI-generated code in Python, TypeScript, and JavaScript sourced from over 1,000
Key concepts
- DECODE
- A dataset containing 53.6K real-world in-IDE code edits of AI-generated code in Python, TypeScript, and JavaScript, sourced from over 1,000 developers.
- In-IDE Edits
- Edits that happen directly within the developer's development environment rather than just final repository data. This captures the temporal sequences of actions a human takes while interacting with AI code suggestions.
- Trajectory-based Supervised Fine-Tuning
- A proposed improvement where models are trained on sequences of code snapshots over time, instead of single points. This helps models learn the progression between different developer states and anticipate future refinements.
Terminology
Summary
This paper introduces DECODE (Developer Edits of Code Dataset), a collection of 53.6K real-world in-IDE code edits of AI-generated code in Python, TypeScript, and JavaScript sourced from over 1,000 developers. It addresses a critical gap in AI programming assistants: while current LLMs are trained on Git data, such data lacks important context such as intermediate edits and temporal ordering of actions.
By providing realistic and granular information on editing behavior,
DECODE enables better alignment between AI-generated suggestions and actual developer workflows.
The DECODE Dataset
DECODE is constructed from a Visual Studio Code extension that tracks localized edits to actual accepted AI completions.
Unlike existing datasets derived from Git commits, which capture relatively polished checkpoints,
DECODE captures the iterative process of modification. The data extraction pipeline follows a four-step process:
-
Obtaining code files by identifying developers who accepted AI completions and tracking file changes.
-
Extracting edits using
git diff --histogramto isolate the specific lines of code relevant to an edit. -
Cleaning and validating trajectories using a CodeBERTScore threshold of 0.68 to ensure relevance to the original completion.
-
Applying PII redaction via the OpenAI Privacy Filter and other established pipelines to protect developer privacy.
Developer Editing Behavior
The researchers used DECODE to analyze when, why, and how AI-generated code is edited,
uncovering several key patterns. Most edits occur very shortly after acceptance, with a steep drop in the total amount of edits 15 minutes after acceptance.
Furthermore, 31% of edit trajectories result in the complete removal of the AI completion. The study identified four specific types of code edits:
-
Customizing code
: Fine-tuning the completion to better align with developer intent (e.g., renaming variables). -
Improving code quality
: Fixing syntax errors or improving readability and comments. -
Changing code functionality
: Modifying the behavior of the code, such as adding new methods or changing logic. -
Removing
: Edits intended to remove the AI completion from the codebase entirely.
Improving LLM Performance
The paper demonstrates that DECODE can be used to improve an LLM's ability to predict developer edits through two tasks: classifying whether a completion will be deleted, unmodified, or modified,
and generating the final edited code.
The results show that fine-tuning on DECODE enables open-source 3B and 7B models to significantly surpass frontier models
on both tasks. For example, fine-tuned models achieved an increase of +0.17 in F1 score for classification and +0.17 in Levenshtein similarity for generation compared to frontier LLMs. Crucially, the authors note that fine-tuning on DECODE does not degrade code generation ability,
with performance on benchmarks like HumanEval and MBPP remaining stable or improving slightly.
Implications for Future AI Assistants
The findings suggest a necessity for developer-centric machine learning approaches
in the development of future coding agents. Because developers often either accept [code] with few modifications or almost completely remove it,
the researchers argue that models must account for code editability
—the ease with which code can be adapted to a specific context. Future work should move beyond simple code correctness to focus on how well a suggestion aligns with developer intent and the sequential and temporal nature of developer edits.
Improvements for AI systems
1. Trajectory-Based Supervised Fine-Tuning (T-SFT)
-
The Improvement: Transition from training models on static code/comment pairs or final Git commits to training on temporal sequences of (y 0, y 1,, y t), where y represents snapshots of code at specific time intervals after an AI completion is accepted.
-
What the system can do: The model will learn the probabilistic transitions between different developer states (e.g., moving from
improving code quality
tochanging functionality
). This enables the AI to anticipate not just what code is correct, but how a human will likely refine it, allowing for moreadaptable
initial suggestions.
2. Predictive Editability Filtering (Pre-computation Loop)
-
The Improvement: Integrate a lightweight classifier trained on D E C O D E data to evaluate the
editability
andretention probability
of a generated snippet before it is presented to the user. -
What the system can do: If the model predicts a high probability of
removal
(the completion being deleted) or low Levenshtein similarity to likely developer-intended modifications, it will trigger an internal re-generation loop. This prevents the presentation ofbrittle
code that is technically correct but lacks alignment with local coding styles or specific intent, thereby reducing developer abandonment rates.
3. Iterative Contextual Prompting (Trajectory-Aware Inference)
-
The Improvement: Modify the inference engine to ingest the sequence of the last k manual edits made by the developer as a structured
edit trajectory
within the prompt prefix, rather than just treating them as static code context. -
What the system can do: The AI will recognize whether a developer is currently in a
customizing
phase (e.g., renaming variables) or afunctionality
phase (e.g., adding logic). By understanding the intent behind the edit pattern, the next suggestion will be contextually aligned with the developer's current workflow stage, significantly reducing the manual effort required to reach a final implementation.
4. Developer-Centric Evaluation Framework
-
The Improvement: Replace or supplement standard
correctness
metrics (like Pass@k on HumanEval) with a multi-dimensional metric suite including Code Retention Ratio (percentage of AI code remaining after 15 minutes), Abandonment Rate, and Edit Distance to Intent. -
What the system can do: This provides a high-fidelity signal for RLHF (Reinforcement Learning from Human Feedback). Instead of rewarding models that produce
correct
but unadaptable code, the reward model will prioritize completions that require minimal Levenshtein distance to be integrated into the developer's specific codebase.
Abstract
Imperfections in AI-generated code require that software developers modify the generated code manually, or by re-prompting an AI programming assistant. Manual code edits provide more realistic and granular information on editing behavior than Git commits, which only contain final successful code snippets. Yet, due to a lack of high-quality, realistic code editing data, LLMs are mostly trained on publicly available Git data (e.g., commits). To address this gap, we introduce DECODE (Developer Edits of Code Dataset), a dataset of 53.6K real-world in-IDE code edits of AI-generated code in Python, TypeScript, and JavaScript, sourced from 1K+ developers. First, we demonstrate the utility of DECODE for data analysis, obtaining insights on when, why, and how AI-generated code is edited. We find that most edits occur within the first 15 minutes after accepting an AI completion, resulting in the removal of AI completions in 31% of edit trajectories. Second, we use DECODE to benchmark the ability of LLMs to predict code edits. We find that finetuning on DECODE enables open-source 3B models to perform code edit prediction tasks significantly better than frontier LLMs. We then discuss implications of this work, emphasizing the necessity of developer-centric machine learning approaches for future AI programming assistants.
Sources
- Program Synthesis with Large Language Models
- SWE-chat: Coding Agent Interactions From Real Users in the Wild
- Qwen3-Coder-Next Technical Report
- Evaluating Large Language Models Trained on Code
- DeepSeek-V3 Technical Report
- The Llama 3 Herd of Models
- Sequence Transduction with Recurrent Neural Networks
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- Qwen2.5-Coder Technical Report
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- The Era of Real-World Human Interaction: RL from User Conversations
- Adam: A Method for Stochastic Optimization
- StarCoder 2 and The Stack v2: The Next Generation
- Next Edit Prediction: Learning to Predict Code Edits from Context and Interaction History
- Understanding and supporting how developers prompt for LLM-powered code editing in practice
- CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
- Modeling Student Learning with 3.8 Million Program Traces
- Code Llama: Open Foundation Models for Code
- Learning Next Action Predictors from Human-Computer Interaction
- OpenAI GPT-5 System Card
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties