Characterizing Model-Native Skills
cs.AI, cs.CL, cs.LG
Submitted: 2026-04-19
Updated: 2026-09-19
Comments: Published as a conference paper at COLM 2026
Code: https://github.com/huggingface/math-verify
License: http://creativecommons.org/licenses/by/4.0/
The gist: Skills are a natural unit for describing what a language model can do and how its behavior can be changed.
Terminology
Abstract
Skills are a natural unit for describing what a language model can do and how its behavior can be changed. However, existing characterizations rely on human-written taxonomies, textual descriptions, or manual profiling pipelines--all external hypotheses about what matters that need not align with the model's internal representations. We argue that when the goal is to intervene on model behavior, skill characterization should be *model-native*: grounded in the model's own representations rather than imposed through external ontologies. We instantiate this view by recovering a compact orthogonal basis from sequence-level activations. The resulting basis is semantically interpretable but need not correspond to any predefined human ontology; instead, it captures axes of behavioral variation that the model itself organizes around. We validate this characterization on reasoning post-training, using the recovered basis for both SFT data selection and inference-time steering. We develop lightweight proxy interventions to identify which directions are most useful for a given model. Across Llama3-8B and Qwen2.5-3B, selecting data along those directions improves Pass@1 by up to 20% on MATH and 41% on AMC, outperforming data selection based on human-characterized skills. Because the basis lives in activation space, the same directions also serve as steering vectors at inference time, improving Pass@8 by up to 4.8% on MATH--an intervention that human-characterized skills cannot support. We further validate the characterization on safety alignment, where selecting adversarial training data for model-native skill coverage rather than textual diversity yields more sample-efficient learning. These results suggest that recovering skills from the model's own representations, rather than imposing them externally, provides a more effective foundation for intervening on model behavior. Codes are open-sourced.
Sources
- Phi-4 Technical Report
- A Theory for Emergence of Complex Skills in Language Models
- Llama-Nemotron: Efficient Reasoning Models
- Training Verifiers to Solve Math Word Problems
- Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
- Adversarial D'ej\`a Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
- Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
- h4rm3l: A language for Composable Jailbreak Attack Synthesis
- The Llama 3 Herd of Models
- OpenThoughts: Data Recipes for Reasoning Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Skill-Targeted Adaptive Training
- Measuring Mathematical Problem Solving With the MATH Dataset
- GPT-4o System Card
- PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
- Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- NaturalThoughts: Selecting and Distilling Reasoning Traces for General Reasoning Tasks
- AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
- AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection