KV-Skill: Forging Expertise in the Model's Native Language
Zhaowei Han, Xiang Zhang, Bing Han, Kai Liu, Danqi Hu, Jie Liu
cs.LG
Submitted: 2026-08-05
Comments: 17 pages, 4 figures, 18 tables. Zhaowei Han and Xiang Zhang contributed equally to this work. Code: https://github.com/shawnzhg/KV-Skill
Code: https://github.com/shawnzhg/KV-Skill
License: http://creativecommons.org/licenses/by/4.0/
The gist: Task knowledge is commonly stored either as text in the prompt or as an update to model weights.
Terminology
Abstract
Task knowledge is commonly stored either as text in the prompt or as an update to model weights. Text is modular but must be interpreted on every use, while weight adaptation makes the resulting capability difficult to load, remove, or share independently. We introduce KV-Skill, a design space of external factorized operators that a frozen language model reads through a lightweight interface. KV-Skill supports two complementary paths. Registration converts an authored text skill into a text-derived operator and trains a shared per-backbone interface. Reward learning develops a compact latent operator directly from task outcomes, with or without an authored skill. Neither path adds positions to the prompt. Across ten benchmarks and four backbones from three model families, converting text to a KV-Skill consistently makes the same procedural knowledge more effective. On Qwen3.5-4B LiveMath, registration reaches 77.2 accuracy, compared with 23.4 for the source text skill, 52.0 for SkillOpt, and 64.5 for SoftSkill. Under matched reward training and parameter budgets, KV-Skill gives the best result in seven of eight matched settings against soft prefixes, prefix tuning, and LoRA. A post-hoc rank analysis further shows that text-derived operators retain nearly all of their benefit with one task-aligned direction per injection layer, while matched random directions fail. Finally, one shared interface retains three independently loadable KV-Skills without measurable forgetting. These results show that task knowledge can be acquired from text or experience, compressed into an external operator, and deployed separately from the backbone. Code is available at: https://github.com/shawnzhg/KV-Skill
Sources
- KV Cache Steering for Controlling Frozen LLMs
- Text-to-LoRA: Instant Transformer Adaption
- SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine
- Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- Latent Cache Flow: Model-to-Model Communication Without Text
- Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
- SoftSkill: Behavioral Compression for Contextual Adaptation
- Steering Language Models With Activation Engineering
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- KBLaM: Knowledge Base augmented Language Model
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
- TextGrad: Automatic "Differentiation" via Text
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks