MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents
cs.AI, cs.CL
Submitted: 2026-09-13
Updated: 2026-09-18
Journal ref: The Pacific Rim International Conference on Artificial Intelligence (PRICAI), 2026
Code: https://github.com/zhangzhenyu13/SummerClaw
License: http://creativecommons.org/licenses/by/4.0/
The gist: Natural language prompts and skills serve as the strategic backbone of LLM-based agents.
Terminology
Abstract
Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent advances in prompt and skill optimization have achieved notable gains, yet all existing methods optimize a single text template---missing the synergy among multiple complementary strategies. We propose MOSCOPT, a text-native, parameter-free algorithm that jointly optimizes a pool of N skills and a gating skill G that dynamically selects K skills per step. To effectively optimize the skills, we build the EditAdam with internally maintained dual states. Through the three-phase interleaved updates with EditAdam, the system monotonically improves without gradient or parameter tuning. Extensive experiments and detailed ablations across 5 benchmarks and 3 target LLMs demonstrate that MOSCOPT consistently outperforms all baselines, and confirm that both the mixture-of-skills architecture with selective activation and the collective evolution with three-phase interleaving are essential to its superior performance. Code is released https://github.com/zhangzhenyu13/SummerClaw/tree/master/summerclaw/agent trainer/algorithms/moscopt.
Sources
- EvoSkill: Automated Skill Discovery for Multi-Agent Systems
- SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches
- Mixture of A Million Experts
- SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- Humanity's Last Exam
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection