Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost
Zixi Huang, Xiheng Wang, Andrew Wang, William Jurayj, Bernal Jiménez Gutiérrez, Daniel Khashabi, Nicholas Andrews
Johns Hopkins University
cs.CL, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: SpeedRunner is a method for programmatic skill learning that reduces agent cost.
Terminology
Summary
SpeedRunner is a method for programmatic skill learning that reduces agent cost. The paper argues that representing skills as executable code is more cost-effective than natural language, as it offloads reasoning to a cheaper Turing machine. SpeedRunner operates in two phases: a wake phase where an actor generates trajectories, and a sleep phase where an inducer, a coding agent, analyzes trajectories and edits a library of executable skills. The inducer uses a code interpreter to surgically analyze trajectories without placing them in context, avoiding context-window bottlenecks and enabling structured cross-trajectory analysis.
The paper evaluates SpeedRunner against ReAct, Online Prompt Optimization (OPO), Agent Skill Induction (ASI), and Voyager on ScienceWorld, BabyAI, and Crafter, using GPT-5.4-mini. Results show SpeedRunner consistently achieves the frontier in learning and cost reduction. Specifically, SpeedRunner is the only method whose cost decreases over training, and the reduction holds across all three benchmarks — most dramatically on BabyAI, where final usage falls to roughly an eighth of the ReAct baseline.
Performance also improves, with SpeedRunner achieving the strongest final performance on all three benchmarks,
including near-perfect performance on BabyAI.
Ablation studies on BabyAI show that removing the code interpreter has the largest efficiency cost, indicating programmatic trace analysis is especially important for compression.
Call-graph analysis reveals SpeedRunner learns deeper and denser call graphs
than baselines, indicating its learned skills are not merely independent wrappers around primitive actions, but are arranged into reusable hierarchical routines.
For example, on Crafter, SpeedRunner reaches maximum depth 8.7 and density 0.149, compared to Voyager's 5.7 and 0.0005.
The paper also demonstrates robustness. Under external randomness in Crafter, varied by zombie spawn frequency, only SpeedRunner improves over the baseline in every condition.
Under distribution shift in ScienceWorld, switching from task 3 to task 4, SpeedRunner achieves the best performance and strongest token compression on both tasks
and loses only 3.3 percentage points on task 3 while reducing inference cost by 27.1%, making it the only method that becomes more efficient on the original task after adapting to a new distribution.
The paper concludes that representing skills as code can achieve the best cost effectiveness while retaining performance,
and that this trend is consistent across distribution shifts, randomness, domains, and models. The key to this success is framing trajectory analysis as an agentic coding task.
Improvements for AI systems
Improvements to AI Systems:
-
Implement a dual-phase skill acquisition pipeline where an actor generates raw interaction trajectories, and a separate coding agent (the inducer) converts those trajectories into executable Python-like skill functions. This offloads reasoning from expensive LLM inference during execution to cheaper code interpretation, reducing per-step token costs.
-
Add a code-interpreter-based trace analyzer that inspects trajectory logs programmatically (e.g., parsing state-action pairs, extracting subgoal boundaries) without feeding full trajectories into the LLM context. This avoids context-window saturation and enables cross-trajectory pattern mining (e.g., detecting repeated action sequences) that would be impossible with sequential in-context analysis.
-
Maintain a hierarchical skill library where newly induced skills are composed into call graphs, not flat wrappers. The system should actively merge overlapping skills, create higher-order functions that call lower-level skills, and prune redundant ones—leading to deeper, denser reuse (e.g., reaching call depth 8.7 vs. 5.7 in baselines).
-
Enable cost-adaptive inference by having the system automatically switch from LLM-driven reasoning to executing pre-compiled skills once a task becomes familiar. The system should track per-task token usage and dynamically reduce reliance on the LLM as skill coverage grows, achieving monotonic cost reduction over training (e.g., dropping to 1/8th of baseline cost).
-
Build robustness to distribution shift and randomness by storing skills as modular, parameterized code (e.g., functions with adjustable parameters for environment variables like zombie spawn rate). When the environment changes, the system should re-induce only affected skills via targeted trajectory analysis, not retrain from scratch—preserving performance on the original task while improving on the new one.
-
Add a structured cross-trajectory compression module that groups similar trajectories, extracts common subroutines, and compiles them into reusable functions. This reduces the number of distinct skills needed and increases the density of the call graph, improving generalization to unseen tasks.
What the Improved AI System Can Do:
-
Solve interactive tasks (e.g., navigation, crafting, instruction following) with near-perfect accuracy while using 8x less inference cost than ReAct-style agents.
-
Automatically discover and reuse hierarchical routines (e.g.,
collect wood
→craft tools
→build shelter
) without human annotation, achieving deeper skill composition than Voyager. -
Maintain high performance under environmental randomness (e.g., varying enemy spawn rates) and distribution shifts (e.g., new task variants) while actually becoming more token-efficient on the original task after adaptation.
-
Scale to long-horizon tasks by never exceeding context limits—trajectory analysis happens offline in a code interpreter, not in the LLM prompt.
-
Adapt to new domains by re-inducing only the necessary skills, with final inference costs that decrease monotonically as the skill library matures.
Abstract
Recently, the practice of augmenting LLM agent capability with skills has gained prevalence. We explore the cost effective adaptation of agents to novel domains by means of learning skills. Existing works focus on performance gain over cost effectiveness. As a result, little is known about what skill learning strategies save cost. We argue that among all the different skill learning methods, those that view skills as programs can achieve the best cost reduction. By executing sequences of actions deterministically, a program-augmented agent can reliably and cheaply achieve goals that would otherwise require trial and error and risk degenerate behavior over long horizons. An agent can learn at inference time by incrementally discovering these programs and equipping them for future tasks. We hypothesize that past trajectories contain enough signal to guide skill learning, even without replay or validation, provided the agent can learn to analyze them. To test our claims, we propose SpeedRunner, a coding agent that analyzes trajectories and refactors skills for better performance on future tasks. Across three different embodied environments, we show that SpeedRunner consistently achieves the frontier in learning and cost reduction while remaining robust against distribution shifts and environmental randomness.
Sources
- A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- Learning without Forgetting
- SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- Automatic Prompt Optimization with "Gradient Descent" and Beam Search
- AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts
- RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
- Continual Learning of Large Language Models: A Comprehensive Survey
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
- ReGAL: Refactoring Programs to Discover Generalizable Abstractions
- SkillX: Automatically Constructing Skill Knowledge Bases for Agents
- Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
- A Comprehensive Survey of Continual Learning: Theory, Method and Application
- Inducing Programmatic Skills for Agentic Tasks
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
- ExpeL: LLM Agents Are Experiential Learners
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering