The Router Within: Eliciting Native Skill Routing from a Frozen LLM
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-14
Updated: 2026-09-28
Code: https://github.com/SWE-agent/mini-swe-agent
License: http://creativecommons.org/licenses/by/4.0/
The gist: Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one.
Terminology
Abstract
Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by preloading every skill's metadata into the context, which disperses the agent's attention and caps the library size. Retrieval pipelines move the selection out of the context, but also out of the agent's capability. We show that the frozen agent LLM already carries the routing signal in its own forward passes, and that two linear maps suffice to read it out with no skill text in the context. Gavel (Glance And Verdict from a frozen LLM) reads it in two steps. A glance projects the task's and each skill's mid-layer states through the two maps, the only parameters trained, and scores the full library against compact per-skill banks that one forward pass builds at installation. A verdict then resumes the shortlisted skills' forward passes and reads the model's own likelihood and yes/no judgment, fused with the glance as a product of experts. Trained once, Gavel transfers zero-shot to three public benchmarks and SkillTraj, our new benchmark of 372 simulated agent trajectories. On Qwen3-32B it outperforms progressive disclosure and retrieve-and-rerank pipelines that add 1.2B to 16B external parameters, by up to 13.4 points on written tasks and up to 21.9 when the need for a skill arises mid-rollout. Routing accuracy improves as the backbone does, and in a bash-agent harness the same 32B triggers the correct skill on Skill-Use more often than far larger frontier models running in Codex.
Sources
- The Scaling Laws of Skills in LLM Agent Systems
- Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees
- MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
- SkillReducer: Optimizing LLM Agent Skills for Token Efficiency
- Gemma 4 Technical Report
- Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?
- SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Agent Skills: A Data-Driven Analysis of Claude Skills for Extending Large Language Model Functionality
- Skill Retrieval Augmentation for Agentic AI
- Representation Learning with Contrastive Predictive Coding
- Skill Is Not Document: Query-Conditioned Compatibility for LLM Agent Skill Routing
- Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
- Qwen3 Technical Report
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- SkillRouter: Skill Routing for LLM Agents at Scale
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks