GitSkills: A Dataset of Agent Skills on GitHub

arXiv:2608.10906 · cs.SE, cs.AI · Submitted 2026-08-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "GitSkills: A Dataset of Agent Skills on GitHub".

Jane: The paper was written by the authors from University College London and University of Hohenheim and University of Cagliari.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Data Collection Mechanics: Jane: So, building on our understanding of agent skills, the second segment zeroes in on the nuts and bolts—how did they actually scrape and process this data from GitHub? It sounds incredibly complex.

Tom: It is! The paper details a whole pipeline. They aren't just looking at the main code files; they are parsing specific metadata, like that `SKILL.md` file, which is essentially acting as a formal description of the skill itself within the repository structure.

Lu: What I found most interesting about the methodology was their approach to documentation enrichment. For each representative skill, they don't just grab the code; they download and parse the front matter from that `SKILL.md` file, which is where they record bundled scripts and reference documents up to a size cap.

Meng: That sounds like a lot of pre-processing overhead just to catalogue *what* was supposed to be there. But collecting repository metadata—star counts, commit histories—that adds an economic layer to the skill definition, doesn't it? It tells us how valuable or popular that skill is perceived to be.

Jane: Exactly! And they are being really meticulous about tracking the lineage of the data. They collect repository metadata *and* they track things like star counts, which gives us a measure of community validation for that skill's existence on GitHub.

Tom: But I want to focus on the historical aspect—the commit history. The paper notes they retrieve the commit history of that `SKILL.md` file, and they even store the author account of the first and last commit, albeit anonymized, which is a deep dive into development provenance.

Lu: The anonymization process itself speaks volumes about ethical data handling in AI research. Masking email addresses and replacing personal names with stable codes ensures that while we can track the *persistence* of an authorship pattern, we aren't compromising individual privacy—that’s crucial for adopting large-scale datasets like this.

Meng: And they mention collecting the commit author accounts and personal names in trailer lines, which are then replaced by codes. That stability across the dataset, as long as those codes can't be reversed, is what allows for longitudinal analysis of authorship trends over time using "GitSkills: A Dataset of Agent Skills on GitHub."

Lalam: This focus on provenance—tracking who contributed and when—is how we teach AI systems not just *what* to do, but *how* development teams actually collaborate. It

Paper discussion segment 2: Tom: So, if I’m hearing you right, Jane, this paper titled "GitSkills: A Dataset of Agent Skills on GitHub" basically solved a massive problem by creating a structured map of human coding knowledge for AI.

Jane: Exactly, Tom. Think of it like this: before this dataset existed, figuring out what an AI could learn from the sprawling mess of GitHub was nearly impossible. This team built a systematic way to catalog and organize those skills.

Lu: And what's revolutionary here isn't just the sheer size of the dataset—which is staggering—but that they’ve applied such rigorous engineering principles to its construction, right? It turns unstructured code into structured knowledge graphs for AI agents.

Meng: From an engineering standpoint, I’m really interested in their enrichment process. They aren't just pulling raw files; they are analyzing the commit history and metadata too. That gives context—it tells us *how* a skill evolved, not just what the final code looks like.

Lalam: That context is everything because it moves AI beyond mere imitation and toward understanding development process itself, which fundamentally improves how we approach human-computer collaboration.

Tom: It’s huge! So, the implication isn't just that AI can write code; it's that AI can now understand the *lifecycle* of software development by learning from millions of real-world examples.

Jane: I think the simplest way to explain the impact is this: instead of us needing a specialist for every niche problem, future AI agents could be trained on these diverse "skills" and dynamically assemble solutions across multiple domains.

Lu: Imagine an AI that doesn't just solve a coding challenge but also knows which version control system was best used, or what the community consensus was on that feature five years ago. That’s the kind of historical wisdom this dataset unlocks for AI.

Meng: But Jane touched on something critical about implementation: if we want to build agents that use this, we need robust APIs that can query not just the code itself, but also the specific metadata they collected—the star counts, the location class... it’s all about the structured layer.

Lalam: The ability to measure and compare skills across different repositories and time periods is what elevates this from a mere archive to a true educational tool for advancing human culture alongside AI.

Tom: It makes you think about how quickly AI will evolve from being just a writing assistant into something that acts like an entire, highly knowledgeable development team.

Jane: We're talking about agents that can really bridge the gap between theory and practice, which is always where the biggest leaps happen in technology.

Lu: And the possibilities for personalized education are wild; instead of generic tutorials, AI could direct someone to a specific GitHub skill perfectly matched to their current weakness!

Meng: Provided we can actually manage the computational load of running those highly complex, context-aware agents on this scale.

Lalam: This dataset isn't just about code; it's about institutionalizing collective human genius, giving AI the ultimate reference library for progress. So next, I wonder how this structured knowledge might impact entirely new forms of creative output...

Paper discussion segment 3: Tom: So, if we’re tracking the progress of *GitSkills: A Dataset of Agent Skills on GitHub*, what does the paper suggest for making these agent skills even better?

Jane: Basically, it shows that just having a list of skills isn't enough; they need standardized ways to interact with those skills so that any system can use them easily.

Lu: Exactly! I think the biggest improvement Lu sees is moving past just *retrieving* a skill and into dynamically *combining* them in novel, complex ways that human developers might not even think of themselves.

Meng: But Lu, if you're combining skills, how do we know which combination is actually stable? You can't just stack functions; you need a reliable orchestration layer that handles dependencies and error states gracefully.

Lalam: Meng brings up a critical point about reliability—the true improvement isn't just in the breadth of skills, but in creating an entirely new trust framework for autonomous code execution across different repositories.

Tom: A trust framework—that’s huge! So we're talking about agents that can self-validate their proposed workflow before touching any code, right?

Jane: That's right; instead of just saying "I can fix this," the agent would have to show its work and prove the fix won't break something else in the process.

Lu: And think about how that shifts development entirely! Instead of spending time debugging poor documentation, we spend time guiding incredibly sophisticated AI agents through complex architectural decisions.

Meng: I agree with Lu, but from an engineering standpoint, this implies a massive overhaul of our testing suites; we'd need new kinds of simulations that test the *interaction* between skills, not just the skills in isolation.

Lalam: It suggests a cultural shift in how we value expertise—the focus moves away from manual execution and toward defining and curating the abstract rules and constraints that guide machine intelligence.

Tom: So, moving beyond the individual skill, we’re building an entire automated ecosystem of reliable knowledge transfer using this dataset?

Jane: It really suggests that AI isn't just a coding assistant anymore; it's becoming a systemic collaborator that understands the whole lifecycle of a project.

Lu: The implications for scientific discovery are staggering; imagine researchers automating the entire process from hypothesis generation to experimental simulation, guided by these combined skills.

Meng: Or think about enterprise software maintenance—instead of having one senior engineer bottlenecking all knowledge, this system could distribute that expertise across hundreds of verifiable, reusable skills.

Lalam: And that scalability has profound social implications; it democratizes complex technical knowledge, allowing small teams or individuals to tackle problems previously reserved only for massive corporate infrastructure.

Tom: Wow, so we're talking about fundamentally changing how we build and maintain all software in the world. But what does this mean for the next generation of developers?

Conclusion: Tom: So, wrapping up our deep dive today, what sticks with me most about this work is how it fundamentally changes what we think of "agent capability."

Jane: Exactly, Tom; it moves beyond just talking *about* skills and gives us a massive library showing how those skills are actually implemented in code across real-world projects.

Lu: I keep thinking about the sheer breadth of the dataset, Jane; it suggests that the next generation of AI won't be built with singular, massive models, but through assembling these highly specific, proven components.

Meng: Building on Lu’s point about assembly, it makes me wonder what the standardization process for these skills looks like—if we can catalog them this thoroughly, can we build a true plug-and-play ecosystem right now?

Lalam: It goes beyond just plugins, Meng; I see it as a shift in culture where the ability to connect existing functionalities becomes more valuable than creating everything from scratch.

Tom: You’re right, Lalam; it’s about composition, which is what I kept circling back to—the sheer power of combining small, verified units of work.

Jane: And that’s fantastic because for developers who aren't AI experts themselves, this dataset provides a roadmap for integration that feels much less daunting.

Lu: Because the authors meticulously cataloged the front matter and even traced commit histories, it adds a layer of verifiable provenance to every single skill listed, which is huge for trust.

Meng: The reliability aspect is what really excites me from an engineering standpoint; knowing where a skill came from and how stable that source repository is cuts down months of architectural guesswork.

Lalam: Considering the collective impact, this work on "GitSkills: A Dataset of Agent Skills on GitHub" doesn't just improve AI; it improves how people collaborate with technology itself by making those interactions more predictable.

Tom: Predictability and scale—those are the two words that really summarize the implications here; it opens up an entire industrial layer of development we hadn't fully accounted for.

Jane: Well, this has been such a fascinating look into how AI tools are becoming modular pieces of software rather than monolithic black boxes.

Lu: I can’t wait to see what kind of wild combinations people start building once they have this foundation to play with!

Meng: Yeah, I'm already thinking about the API wrappers we'd need just to manage the querying across these different source types.

Lalam: We really appreciate you walking us through the depth of "GitSkills: A Dataset of Agent Skills on GitHub" today; it’s going to change how people think about software construction.

Tom: Thanks to all of you for joining us; we'll take a quick break and when we come back, we're going to shift gears and look at some recent work in multimodal reasoning...

University College London · University of Hohenheim · University of Cagliari

cs.SE, cs.AI

Submitted: 2026-08-11

Updated: 2026-09-11

Comments: Giuseppe Destefanis, Daniel Graziotin, Matteo Vaccargiu, and Marco Ortu. 2027. GitSkills: A Dataset of Agent Skills on GitHub. In Proceedings of the 24th International Conference on Mining Software Repositories (MSR '27). Association for Computing Machinery, New York, NY, USA, 3 pages. To appear

Code: https://github.com/giuseppedestefanis/gitskills-sample

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: GitSkills is a dataset of 3,797,117 SKILL.md files collected from 282,200 public GitHub repositories owned by 195,841 accounts, gathered in July 2026.

Key concepts

GitSkills: A Dataset of Agent Skills on GitHub
This dataset systematically catalogs human coding knowledge found across real-world GitHub projects. It aims to provide a structured map of skills, allowing AI agents to learn and assemble solutions from diverse, verifiable components.
Data Provenance/Commit History
This refers to tracking the development history of a skill's source file, including who wrote it and when. The dataset records the first and last commit authors, providing deep insights into how a skill evolved over time.
Metadata Enrichment
This process involves collecting context beyond just the raw code files. It includes gathering repository metadata (like star counts) and parsing specific files to give a measure of the skill's popularity and scope.
Systemic Collaboration
This describes AI agents moving beyond acting as simple writing assistants. Instead, they become systemic collaborators that can dynamically combine multiple verified skills to solve complex problems across an entire project lifecycle.

Terminology

Summary

GitSkills is a dataset of 3,797,117 SKILL.md files collected from 282,200 public GitHub repositories owned by 195,841 accounts, gathered in July 2026. An agent skill is defined as a folder containing a SKILL.md file with instructions for a language-model agent, optionally accompanied by scripts and reference files, a format introduced by Anthropic in October 2025 as an open specification. The dataset groups identical files by content hash into 1,877,981 distinct contents, retaining every file occurrence with its repository, path, and content hash, and enriches one representative per group with full text, parsed front matter, folder contents, repository metadata, and, for a subset, the commit history of the file.

The dataset is stored as a single self-contained SQLite file with tables including artifacts (3,797,117 rows, one per discovered file with repository, path, basename, location class, content hash, representative flag, and for representatives full text, front matter, and body size), repos (282,200 rows with owner, star count, primary language, fork status, creation date, last-push date), artifact siblings (7,264,865 rows for files stored alongside representative skills, including path, entry type, size, and text of files under a size cap), and mining runs (7 rows with query, timestamps, and result counts). Commit history, covering first and last commit dates, anonymized author accounts (user or bot), and commit count, is available for 458,548 skills in standard locations and a size-stratified sample of the rest.

The collection pipeline is read-only, using GitHub code-search and REST APIs. Because code search returns at most 1,000 results per query and its reported totals proved unreliable (roughly 349,000 for the filename query against over 3.8M files retrieved), discovery partitions the search space by file size until every range can be retrieved completely. Files are grouped by content hash; one representative per group is enriched, preferring a file in the.claude/skills/ directory, and all copies are retained. The dataset covers public repositories only and should be read as a lower bound on the population, since GitHub code search indexes only default branches, files under 384 KB, recently active repositories with fewer than 500,000 files, and forks only when they have more stars than the parent repository.

The paper notes that skills are unlike typical mined artifacts: "they are written mainly in natural language, a model selects them probabilistically at run time, and no compiler or type checker verifies the selection. They also have no central registry or package manager, so they spread by copying folders between repositories." Reuse is significant, with 50.5% of collected files being verbatim copies. The dataset supports research on adoption, reuse, structure, authorship, maintenance, and security of agent skills, with proposed research questions covering adoption and linguistic evolution, development of a shared format, reuse without a package manager, software metrics for natural-language instructions, and maintenance and trust (including supply-chain attack analogs and outdated skills).

Anonymization replaces commit author accounts with keyed one-way codes, identical for the same account throughout, while bot accounts keep their login; email addresses and personal names in commit messages are redacted, and AI assistant names in Co-authored-by trailers are kept. The dataset is archived on Zenodo, with a Parquet mirror on Hugging Face and a sample on GitHub.

Improvements for AI systems

Improvements to AI Systems:

  1. Skill-Aware Retrieval and Ranking: Train a retrieval model on the 1.88M distinct skill contents to embed natural-language instructions and match them to user tasks. The improved system can, given a user query, retrieve the most relevant SKILL.md files from the dataset, rank them by semantic similarity, and suggest reusable skills—even across repositories—without requiring a package manager.

  2. Skill Composition and Adaptation: Use the 50.5% verbatim-copy rate and repository metadata (stars, language, fork status) to train a model that predicts which skill variants are most effective for a given context. The improved system can automatically adapt a retrieved skill by adjusting its instructions to match the target repository's language, file structure, and maintenance patterns, reducing the need for manual rewriting.

  3. Supply-Chain Risk Detection for Agent Skills: Leverage the commit history and anonymized author codes to train a classifier that flags skills with anomalous maintenance patterns (e.g., sudden large edits, bot-only commits, or copying from unvetted sources). The improved system can, before executing a skill, assess its trustworthiness and warn users about potential supply-chain attacks or outdated instructions.

  4. Natural-Language Instruction Quality Scoring: Using the front matter, body size, and sibling file structure, train a model to predict the clarity, completeness, and actionability of a skill. The improved system can automatically score new or user-written skills, suggest missing sections (e.g., inputs, outputs, error handling), and flag ambiguous or overly verbose instructions—acting as a linter for natural-language agent instructions.

  5. Cross-Repository Skill Generalization: Train a model on the artifact siblings table (7.26M files) to learn which auxiliary scripts and reference files are commonly paired with skills. The improved system can, given a new skill, automatically generate or recommend missing companion files (e.g., a test script or config template) based on patterns from similar skills in the dataset.

  6. Skill Lifecycle Prediction: Using first/last commit dates and commit counts, train a time-series model to predict whether a skill is actively maintained, abandoned, or likely to be deprecated. The improved system can proactively suggest updates or replacements for outdated skills in a user's repository, based on the dataset's observed maintenance trajectories.

  7. Anonymized Author Behavior Modeling: Using the keyed one-way author codes, train a model to identify patterns of skill authorship (e.g., solo vs. collaborative, bot-assisted vs. human). The improved system can infer the likely reliability and style of a skill's author, enabling personalized recommendations (e.g., preferring skills from authors with a history of long-lived, well-maintained skills).

Sources

Related papers