AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficient Kernel Generation
cs.CL, cs.AI, cs.LG
Submitted: 2026-04-17
Updated: 2026-09-05
Comments: Preliminary work. The implementation is available at https://github.com/StigLidu/AdaExplore
Code: https://github.com/StigLidu/AdaExplore
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent large language model (LLM) agents have shown promise in using execution feedback for test-time adaptation.
Terminology
Abstract
Recent large language model (LLM) agents have shown promise in using execution feedback for test-time adaptation. However, robust self-improvement remains far from solved: most approaches still treat each problem instance independently, without accumulating reusable knowledge. This limitation is particularly pronounced in domain-specific languages such as Triton, which are underrepresented in LLM pretraining data. Their strict constraints and non-linear optimization landscape further make naive generation and local refinement unreliable. We propose AdaExplore, an agent framework that enables self-improvement via accumulated execution feedback for performance-critical kernel code generation through two complementary stages: failure-driven adaptation and diversity-preserving search, jointly improving correctness and optimization performance without additional fine-tuning or external knowledge. In the adaptation stage, the agent synthesizes tasks and converts recurring failures into a reusable memory of validity rules, helping subsequent generations remain within the feasible set. In the search stage, the agent organizes candidate kernels as a tree and alternates between small local refinements and larger structural regeneration, allowing it to explore the optimization landscape beyond local optima. Experiments on kernel runtime optimization benchmarks validate these gains: AdaExplore achieves 3.11x and 1.62x speedups on KernelBench Level-2 and Level-3, respectively, within 100 steps, and continues to improve with additional computation.
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization
- Kevin: Multi-Turn RL for Generating CUDA Kernels
- Qwen3-Coder-Next Technical Report
- Evaluating Large Language Models Trained on Code
- Exploring and Controlling Diversity in LLM-Agent Conversation
- CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
- Qwen2.5-Coder Technical Report
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
- CoCoEvo: Co-Evolution of Programs and Test Cases to Enhance Code Generation
- AutoTriton: Automatic Triton Programming with Reinforcement Learning in LLMs
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
- Scattered Forest Search: Smarter Code Space Exploration with LLMs
- Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- KernelBench: Can LLMs Write Efficient GPU Kernels?
- HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
- Code Llama: Open Foundation Models for Code
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering