Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization via an Experience-Driven Workflow and Experience Graph Memory
cs.LG, cs.MA
Submitted: 2026-08-26
Updated: 2026-08-26
Terminology
Sources
- Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction
- Evaluating Large Language Models Trained on Code
- From Tokens to Regions: CUDA-Sensitive Instruction Tuning for GPU Kernel Generation
- CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
- CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- ChipNeMo: Domain-Adapted LLMs for Chip Design
- MemGPT: Towards LLMs as Operating Systems
- Fine-Tuning GPT-5 for GPU Kernel Generation
- Astra: A Multi-Agent System for GPU Kernel Performance Optimization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks