CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language
cs.CL, cs.AI, cs.MA, cs.PL, cs.SE
Submitted: 2026-08-30
Updated: 2026-08-30
Code: https://github.com/anomalyco/opencode
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation, correctness validation, and hardware-aware parallel optimization, creating a substantial expertise
Terminology
Abstract
Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation, correctness validation, and hardware-aware parallel optimization, creating a substantial expertise barrier and making generating CUDA kernels directly from natural language (Text2CUDA) essential. Meanwhile, the general-purpose code generation capability of Large Language Models (LLMs) prompts a series of works exploring LLM-based CUDA kernel generation. They mainly focus on transpilation from high-level frameworks such as PyTorch to CUDA (Torch2CUDA) rather than Text2CUDA, where models must understand the high-level input semantics and handle low-level kernel implementation and validation. Additionally, these methods are vulnerable to reward hacking due to reliance on predefined test inputs. In this paper, we propose CUDA-Harness, a framework for harnessing agentic CUDA kernel generation and optimization from natural language. Specifically, we introduce Intermediate-Structured Generation to connect high-level semantic understanding with low-level kernel generation. To dilute reward hacking in Text2CUDA, we construct Synthesis-Based Verification to provide isolated test data and progressive validation. Furthermore, we propose Feedback-Adaptive Evolution, a kernel evolution strategy that prioritizes correctness while optimizing performance. Finally, through extensive experiments, we demonstrate the effectiveness of CUDA-Harness, with further evaluations illustrating generalization across LLMs, hardware platforms, and to C-to-CUDA transpilation.
Sources
- cuPilot: A Strategy-Coordinated Multi-agent Framework for CUDA Kernel Evolution
- CUDA-LLM: LLMs Can Write Efficient CUDA Kernels
- CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- STARK: Strategic Team of Agents for Refining Kernels
- Kernel-Smith: A Unified Recipe for Evolutionary Kernel Optimization
- From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph
- QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation
- ConCuR: Conciseness Makes State-of-the-Art Kernel Generation
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- KernelBench: Can LLMs Write Efficient GPU Kernels?
- KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization
- ReAct: Synergizing Reasoning and Acting in Language Models
- Towards Automated Kernel Generation in the Era of LLMs
- GLM-5: from Vibe Coding to Agentic Engineering
- CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
- LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
- CUDABench: Benchmarking LLMs for Text-to-CUDA Generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering