Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
cs.AI, cs.CL
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: 48 pages, 3 figures
Code: https://github.com/VectorSpaceLab/AREX-Skill
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Autonomous agents are beginning to carry out machine-learning (ML) research end to end.
Terminology
Abstract
Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how outside the agent. We call this missing layer operational knowledge, the know-how that separates knowing a method from making it work. That knowledge is not absent from the field. It appears in repositories and papers, but in forms written for human readers and too large to load during a task. Once distilled into compact, verified skills, this knowledge can be reused across tasks rather than rediscovered during each run. We present DisCo, a skill-powered research agent that creates skills and uses them during research. Its distillation runs in two complementary forms: task-agnostic, condensing the field's widely used repositories into reusable skills, and task-oriented, producing the skills a concrete task calls for. The former, applied across the open ecosystem, yields the AREX-Skill Library, with 5,000+ verified skills distilled from 1,000 widely used ML repositories and organized into 20 areas and 178 capability families. With the GPT-5.5 backbone, research harness, and downstream execution budget held fixed, the skill-equipped research agent scores 134.3% higher on MLE-bench, 34.4% higher on PaperBench, 9.2% higher on FrontierCS, and 14.0% higher on PassNet than the same agent without skills. These gains come from adding distilled operating context under that fixed setup.
Sources
- Toward Autonomous Long-Horizon Engineering for ML Research
- BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
- MARS: Modular Agent with Reflective Search for Automated AI Research
- SELA: Tree-Search Enhanced LLM Agents for Automated Machine Learning
- MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery
- From Anatomy to Smells: An Empirical Study of SKILL.md in Agent Skills
- AIDE: AI-Driven Exploration in the Space of Code
- Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
- The FM Agent
- AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions
- DeepCode: Open Agentic Coding
- PassNet: Scaling Large Language Models for Graph Compiler Pass Generation
- ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- FrontierCS: Evolving Challenges for Evolving Intelligence
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- R&D-Agent: An LLM-Agent Framework Towards Autonomous Data Science
- AIBuildAI: An AI Agent for Automatically Building AI Models
- Reflective Paper-to-Code Reproduction Enabled by Fine-Grained Verification
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection