The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot".
Jane: The paper was written by Fangchen Song, Ashish Agarwal and Wen Wen from University of Texas at Austin.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everybody. Today we are digging into a paper that's been making waves in the software world, and it's called "The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot." Jane, I have to say, this title alone got me excited because it's not just about whether AI writes code faster. It's about how AI changes the way people work together.
Jane: Absolutely, Tom. And I love that they're looking at open-source software specifically. That's where volunteers from all over the world come together to build things like Linux or the web servers that run half the internet. There's no boss telling anyone what to do, so it's a really pure test of how AI changes voluntary collaboration.
Tom: Right, and the authors — Fangchen Song, Ashish Agarwal, and Wen Wen from UT Austin — they had access to proprietary data from GitHub itself. That's huge. They could see exactly which projects were using Copilot, which is GitHub's AI pair programmer, and which ones weren't.
Jane: And that data matters because a lot of earlier studies just looked at individual developers in a lab setting. You know, "here's a task, finish it with or without AI." But this paper looks at real projects, real teams, real messy collaboration. It's like studying how a new power tool changes a construction crew, not just how fast one carpenter can hammer a nail.
Tom: Exactly. And the headline finding is that Copilot increased project-level code contributions by about five point nine percent. But here's the twist — it also increased coordination time by about eight percent. So people are writing more code, but it's taking longer to get that code integrated into the project. That's a tradeoff we don't usually hear about.
Jane: And that tradeoff is the heart of the paper. It's not just "AI makes everything faster." It's "AI changes the whole rhythm of collaboration." We'll get into the details in a bit, but I think the big takeaway for listeners is that this study is using real-world data to show that generative AI doesn't just affect individual productivity — it affects how teams coordinate, how discussions happen, and even who ends up contributing more.
Tom: And that's what makes this paper so important. It's not a toy experiment. It's a real look at how AI is reshaping the open-source ecosystem, which is worth trillions of dollars to the global economy. So stay tuned, because we're going to break down the methodology and the surprising findings about core versus peripheral developers next.
Summary: Jane: So, Tom, we've set the stage. Now let's talk about what the paper actually found in more detail. The researchers used something called the Generalized Synthetic Control Method, which sounds complicated, but the idea is pretty simple. They built a "synthetic twin" for each project that used Copilot, using projects that didn't use it, to estimate what would have happened without the AI.
Tom: That's a clever way to handle the fact that projects adopt Copilot at different times. And what they found is that Copilot didn't just increase the total number of merged pull requests — that's the five point nine percent — but it also brought in more developers. Developer coding participation went up by three point four percent, and individual productivity went up by two point one percent.
Jane: So more people are showing up, and the people who are already there are producing more. But then, the coordination time — the time it takes from when someone submits a pull request to when it actually gets merged into the project — went up by eight percent. And the reason seems to be that there's just more discussion happening. Comments per merged pull request went up by six point five percent.
Tom: And that's fascinating because it means AI isn't replacing human communication. It's actually generating more of it. The paper suggests that AI helps people understand algorithms and explore alternative solutions, so developers have more to say. Plus, new developers who might have been too intimidated to contribute before are now jumping in, and their code needs more clarification.
Jane: Right, and that brings us to one of the most interesting parts — the difference between core and peripheral developers. Core developers are the ones who run the project, design the architecture, and have deep knowledge of the codebase. Peripheral developers are more like occasional contributors who fix bugs or add small features.
Tom: And the paper found that core developers benefited more. Their share of code contributions went up by about six point five percent, while peripheral developers saw a bigger increase in coordination time. That's a really important finding because it suggests AI might be widening the gap between the people at the center of a project and the people on the edges.
Jane: And the reason, the authors argue, is project familiarity. Core developers know the project so well that they can use AI to generate ideas that actually fit the architecture. Peripheral developers, on the other hand, might generate code that looks good but doesn't quite fit the project's conventions, so it needs more discussion and revision.
Tom: So the summary is that AI helps everyone, but it helps the people who already know what they're doing even more. And that has big implications for who gets credit, who gets visibility, and who sticks around in open-source communities. We'll dig into those implications next.
Improvements and Implications: Tom: So, Jane, we've covered the findings. Now let's talk about what this means for the future. The paper doesn't just stop at "AI increases contributions and coordination time." It also suggests some practical improvements for how open-source communities should respond.
Jane: Right, and one of the big suggestions is governance. If AI is making core developers more productive but also making it harder for peripheral developers to get their code merged in a timely way, then communities need to step in. The paper suggests things like structured pull-request templates, mentoring programs, and even protocols for validating AI-generated code.
Tom: That's a really practical takeaway. Because if peripheral developers feel like their contributions are just sitting there, waiting for review, they're going to get discouraged and leave. And that would be a real loss, because peripheral developers bring fresh perspectives and catch bugs that core developers might miss.
Jane: Exactly. And there's a bigger implication here for companies too. The paper argues that in corporate software teams, the people who are most central to decision-making and system design are likely to benefit more from AI pair programmers. So if you're a manager, you might want to think about how to support the developers who are doing more routine work, so they don't get left behind.
Tom: And that connects to something the paper says about coordination becoming the new bottleneck. Before AI, the bottleneck was writing code. Now that AI can write code faster, the bottleneck is integrating it, reviewing it, and making sure it all fits together. That's a huge shift in how we think about software development.
Jane: And the paper even looks at code quality. They found that while the total number of issues and bugs went up, the number of issues per pull request stayed about the same. So the code quality isn't getting worse — there's just more code, so naturally more bugs. That's reassuring.
Tom: But here's the thing that really got me thinking. The paper suggests that AI might be shifting the balance of power in open-source communities. Core developers are getting more productive, and peripheral developers are facing more friction. Over time, that could mean fewer new contributors, less diversity of thought, and maybe even a more concentrated group of people controlling the most important projects.
Jane: And that's a real concern, because open-source software thrives on broad participation. The paper doesn't say this is inevitable, but it does say communities need to be proactive. They need to think about how to make AI work for everyone, not just the people at the top.
Tom: So the improvements the paper suggests aren't just technical — they're social and organizational. It's about designing the community so that AI amplifies collaboration instead of concentrating it. And that's a really forward-thinking way to look at this. I'm excited to see how these ideas get picked up by real projects.
Conclusion: Tom: Alright, Jane, we've covered a lot of ground on "The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot." Let's wrap this up.
Jane: Yeah, let's do it. The big picture is that Copilot increased project-level code contributions by about five point nine percent, but it also increased coordination time by eight percent. And the reason for that is more discussion, more people participating, and more back-and-forth before code gets merged.
Tom: And we saw that core developers benefited more than peripheral developers. Their share of contributions went up, while peripheral developers faced longer coordination times. That's a warning sign that AI might be concentrating influence in open-source communities.
Jane: But the paper also gives us hope. It shows that AI can bring more people into the fold, as long as communities adapt. Structured templates, mentoring, and better review processes can help peripheral developers use AI effectively and get their contributions merged faster.
Tom: And for companies, the lesson is that coordination is becoming the new bottleneck. If you want to get the most out of AI, you need to invest in the processes that help people work together, not just the tools that help them write code.
Jane: Exactly. And I think that's the lasting contribution of this paper. It moves the conversation from "does AI make developers faster?" to "how does AI change the way teams work?" That's a much more important question.
Tom: Well said, Jane. So we're going to say goodbye to this paper and get ready for the next one. Thanks to everyone who tuned in, and we'll see you next time.
Jane: Take care, everybody. Keep coding, keep collaborating, and keep asking the big questions.
Fangchen Song, Ashish Agarwal, Wen Wen
University of Texas at Austin
cs.SE, cs.AI, cs.HC, econ.GN, q-fin.EC
Submitted: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 69/100
Key concepts
- GitHub Copilot
- GitHub Copilot is GitHub's AI pair programmer used in software development. The paper analyzed which projects used it and which did not to study its real-world impact on collaboration.
- Coordination Time
- This refers to the time it takes from when a developer submits a pull request until that code is actually merged into the project. The study found this time increased by 8% when using Copilot.
- Core vs. Peripheral Developers
- Core developers are those who run projects, design architecture, and have deep codebase knowledge. Peripheral developers are occasional contributors who fix bugs or add small features; the paper found core developers benefited more from AI.
Terminology
Summary
Summary
This paper investigates the impact of GitHub Copilot, a generative AI pair programmer, on collaborative open-source software (OSS) development. The study addresses three research questions: (RQ1) How do AI pair programmers affect project-level code contributions in OSS development? (RQ2) How do AI pair programmers affect coordination time in OSS development? (RQ3) How do AI pair programmers affect code contributions and coordination time for core versus peripheral developers differently?
The authors use a combination of publicly available data on GitHub projects and proprietary data on Copilot use provided by GitHub. The unit of analysis is at the project-month level, with a sample period from January 2021 to December 2022. The treatment group consists of projects where Copilot was both supported by local coding environments and used by developers to code. The control group includes projects where Copilot was not used throughout the sample period. The final sample includes 7,637 projects, with 4,491 in the treatment group and 3,146 in the control group. The authors estimate their model using the Generalized Synthetic Control Method (GSCM).
The empirical results show that the use of GitHub Copilot is associated with a 5.9% increase in the number of project-level code contributions but also an 8% increase in coordination time.
These findings indicate a tradeoff between contribution gains and coordination time in OSS development following the use of Copilot.
Mechanism analyses reveal that the increase in project-level code contributions is accompanied by a significant increase in both developer participation and individual productivity. Specifically, Copilot use is associated with a 3.4% increase in developer coding participation
and a 2.1% increase in individual code contributions.
At the same time, the increase in coordination time is accompanied by a higher volume of discussions surrounding code contributions, a broader set of developers participating in these discussions, and greater discussion intensity per developer. The authors find a 6.5% increase in code discussion volume per merged PR,
a significant 1.7% increase in code discussion participation,
and a 5.9% increase in code discussion intensity per developer.
Importantly, the combined effect of these two competing forces still yields an overall positive effect on project-level timely merge of code contributions into the codebase. The results indicate that Copilot use is associated with a 3.5% increase in PRs merged within one day, a 4.1% increase in PRs merged within three days, and a 5.1% increase in PRs merged within ten days.
The authors also find heterogeneous effects across developer roles. Core developers experienced a relatively larger increase in project-level code contributions and peripheral developers experienced a larger increase in coordination time.
Specifically, the proportion of merged PRs from core developers significantly increased by 0.054 or 6.5%,
and peripheral developers experienced an increase of 0.085 or 5.4% in their relative average merge time.
Furthermore, "peripheral developers contribute a smaller proportion of timely merge of code contributions, suggesting that the project-level productivity gain could be relatively less for peripheral developers than for core developers."
The study contributes to literature on generative AI in OSS development, generative AI in team collaboration, and heterogeneity in the impact of generative AI among individuals. The authors note that "while the use of AI pair programmers can increase developer participation, it may shift the balance of contributions toward core developers, highlighting the need for governance structures that support the participation of peripheral contributors. The findings also have implications for software development teams in firms, as
developers who are more central to decision-making and system design are likely to benefit more from AI pair programmers and
coordination and integration may become more important constraints in software development, making organizational processes and team structures critical determinants of the returns to generative AI adoption."
Improvements for AI systems
Based on the scientific paper, here are the specific improvements that can be made to AI systems:
- Context-Aware Code Generation for Project Familiarity
-
Current limitation: AI generates code that often lacks project-specific architectural and strategic context, particularly for peripheral developers with limited project familiarity.
-
Improvement: Train AI systems to incorporate project-level context (codebase architecture, historical evolution, design conventions, and future direction) into code suggestions. The AI should analyze the project's existing patterns, dependencies, and coding standards before generating code.
-
Resulting capability: The AI can produce code that aligns more readily with existing project structures, reducing the need for extensive clarification and discussion during code integration.
- Role-Aware Assistance
-
Current limitation: The same AI assistance benefits core developers (who have deep project familiarity) more than peripheral developers, leading to uneven contribution gains and longer coordination time for peripheral developers' code.
-
Improvement: Develop AI systems that detect whether the user is a core or peripheral developer based on their contribution history, project tenure, and familiarity level. For peripheral developers, the AI should provide additional context-rich prompts, explain project conventions, and generate code that better fits the project's existing architecture.
-
Resulting capability: The AI can level the playing field by providing targeted support to less familiar contributors, potentially reducing coordination time and increasing their timely code contributions.
- Discussion Facilitation and Coordination Support
-
Current limitation: AI pair programmers increase code discussions and coordination time by introducing broader perspectives and attracting less familiar contributors, making consensus harder to reach.
-
Improvement: Integrate coordination-aware features into AI pair programmers, such as:
-
Generating code that anticipates potential integration conflicts and proactively suggests alignment with existing code patterns.
-
Providing automatic summaries of code changes in plain language to reduce the need for clarification discussions.
-
Suggesting standardized pull request templates and code documentation that reduce ambiguity for reviewers.
-
Resulting capability: The AI can reduce the volume and diversity of discussion needed to integrate contributions, thereby shortening coordination time while maintaining contribution quality.
- Ideation Enhancement with Contextual Validation
-
Current limitation: AI enhances ideation capabilities but does not validate whether generated ideas align with the project's specific goals, constraints, and history.
-
Improvement: Train AI systems to evaluate generated code suggestions against the project's stated objectives, existing dependencies, and long-term roadmap. The AI should flag suggestions that may conflict with the project's architectural direction.
-
Resulting capability: The AI can help core developers pursue larger, more innovative changes with confidence while helping peripheral developers avoid proposing changes that require extensive rework.
- Familiarity-Adaptive Prompt Engineering
-
Current limitation: Peripheral developers struggle to provide project-specific information necessary to fully leverage AI tools, resulting in code that requires more discussion.
-
Improvement: Develop AI systems that automatically infer project context from the developer's recent activity, the files they are editing, and the project's documentation. The AI should proactively ask clarifying questions about project-specific constraints when it detects low familiarity.
-
Resulting capability: The AI can generate code that better incorporates project-specific context even when the developer cannot articulate it, reducing the need for post-hoc clarification.
- Coordination-Time Prediction and Optimization
-
Current limitation: AI systems do not account for the downstream coordination costs of generated content.
-
Improvement: Train AI systems to estimate the coordination time required to integrate their outputs into collaborative projects. The AI should optimize for a balance between contribution volume and integration efficiency.
-
Resulting capability: The AI can suggest modifications that minimize discussion requirements while maintaining quality, helping teams achieve faster integration.
- Role-Specific Benefit Maximization
-
Current limitation: AI systems do not differentiate between users who derive high expected benefits (core contributors) versus those with lower expected benefits (peripheral contributors).
-
Improvement: Develop AI systems that adjust their output based on the user's role and expected benefits. For core developers, the AI should support experimentation and novel solutions. For peripheral developers, the AI should focus on reducing participation costs and generating code that requires minimal discussion.
-
Resulting capability: The AI can maximize net benefits for all contributors, potentially increasing overall participation and contribution quality.
- Project-Familiarity Scoring Module
-
Improvement: Add a module that computes a familiarity score for each developer-project pair based on contribution history, tenure, and activity patterns. Use this score to adjust the AI's output style and the amount of context provided.
-
Resulting capability: The AI can automatically tailor its assistance to match the developer's level of project knowledge.
- Discussion-Pattern Analysis
-
Improvement: Train AI systems to analyze historical discussion patterns in a project to identify common sources of confusion (e.g., unclear design intent, missing context). Use this analysis to preemptively address these issues in generated code and accompanying documentation.
-
Resulting capability: The AI can reduce the volume of clarification discussions by proactively addressing known pain points.
- Integration-Risk Flagging
-
Improvement: Implement a feature that flags generated code with high integration risk (e.g., code that touches multiple interdependent modules, code that may conflict with existing architectural decisions). The AI should suggest safer alternatives or provide additional context to facilitate review.
-
Resulting capability: The AI can reduce coordination time by identifying and mitigating potential integration issues before they reach the review stage.
These improvements would enable AI systems to not only increase code contribution volume but also reduce the coordination burden, making the overall development process more efficient and equitable across developer roles.
Sources
- Generative AI Enhances Team Performance and Reduces Need for Traditional Teams
- The Impact of AI on Developer Productivity: Evidence from GitHub Copilot
- The Impact of Large Language Models on Open-source Innovation: Evidence from GitHub Copilot
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties