SPT: Skills as Pre-Training Data for Agentic Language Models
cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
Code: https://github.com/huggingface/trl
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agentic (tool-using) language models are mainly trained on tool-call traces and agent trajectories during post-training.
Terminology
Abstract
Agentic (tool-using) language models are mainly trained on tool-call traces and agent trajectories during post-training. These data provide direct behavioral supervision, but producing them requires task environments, execution, and verification, making broad tool and task coverage expensive. Publicly available skills offer another source of training data: they encode reusable tool semantics and workflows but are typically used only as inference-time context. We introduce Skill Pre-Training (SPT), a mid-training method that applies causal language modeling to SkillCorpus, a collection of public multi-file skill packages, optionally mixed with general data. To preserve relations among files within each package, we also introduce Reference Insert, a reference-aware assembly strategy that places supporting files near their mentions in the primary instruction. Experiments across multiple model scales and post-training recipes show that SPT consistently improves agentic performance over mid-training on general or trajectory data, while largely preserving general performance. Data mixture experiments show additional benefits from combining skill data with general annealing corpora. These results indicate that skill packages are a valuable data source for pre-training agentic language models.
Sources
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Efficient Training of Language Models to Fill in the Middle
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Agentic Reinforced Policy Optimization
- Towards General Agentic Intelligence via Environment Scaling
- What Do Agents Learn from Trajectory-SFT: Semantics or Interfaces?
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- Measuring Massive Multitask Language Understanding
- Training Compute-Optimal Large Language Models
- Agentic Tool Use in Large Language Models
- Qwen2.5-Coder Technical Report
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
- Scaling Laws for Neural Language Models
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents
- SkillNet: Create, Evaluate, and Connect AI Skills
- Instella: Fully Open Language Models with Stellar Performance
- StarCoder 2 and The Stack v2: The Next Generation
- 2 OLMo 2 Furious
- APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering