StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning".
Jane: StitchCUDA proposes a multi-agent framework integrated with rubric-based agentic reinforcement learning to automate end-to-end GPU program generation, addressing limitations in prior work that focused only on single kernels.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, to recap, we’re talking about "StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning." The authors are Li, Zhang, Chen, Luo, and Hong. They’re proposing this multi-agent setup to solve the issue where existing methods only focus on optimizing individual kernels instead of the whole program.
Jane: Right. And their core idea is to use this multi-agent system—a Planner for design, a Coder for writing code, and a Verifier for checking correctness and performance—to build complete GPU programs automatically. It’s about orchestrating the whole thing rather than just tackling one piece at a time.
Lu: I found it interesting how they structured the agents to handle system-level design alongside kernel implementation, which suggests they are thinking about data movement and host settings in a holistic way.
Meng: That holistic view is what concerns me practically; if the Planner designs something that’s fundamentally flawed from a data flow perspective, does the Coder just dutifully implement it, or can it catch those big system bottlenecks?
Lalam: The idea of an end-to-end framework means we aren't just getting a snippet of code; we are getting a fully functional piece of software that runs on the GPU. That level of automation is something I think could really streamline our entire development pipeline for AI applications.
The paper's summary: Tom: Now, let’s look at the summary of what StitchCUDA actually does. Basically, it takes a reference PyTorch code and uses this iterative "plan–code–profile–refine" loop to create a working GPU program that performs well.
Jane: The core mechanism they use is this loop where the agents talk to each other in stages: the Planner figures out what needs to be done, the Coder writes the code for a part, and then the Verifier profiles it and sends feedback back so the Coder can improve.
Lu: What really caught my eye was how they tackled making this agentic process better by using rubric-based reinforcement learning over two specific skills: generating from scratch and then optimizing based on feedback.
Meng: That part about decomposing the RL into atomic skills sounds clever for reducing the computational load, but I'm still wondering if that decomposition actually helps the Coder learn more complex optimization strategies than a standard prompt might achieve.
Lalam: I think the rubric-based reinforcement learning is where it gets really smart because they designed rewards to stop it from just guessing or doing something technically correct but useless, which is what we call reward hacking.
The paper's improvements: Tom: Let’s talk about what they claim StitchCUDA actually improved over prior work. They show that this framework achieves a nearly one hundred percent success rate on KernelBench Level three tasks, which is a big jump from earlier attempts like CUDAForge that struggled there <ref:2603.02637#pg2>.
Jane: The authors highlight that StitchCUDA doesn't just write code; it manages cross-kernel optimization and host-side orchestration, which was a major gap in previous single-kernel focused tools. This suggests it can handle the complex system-level requirements of larger models.
Lu: They specifically point out that by using rubric rewards, they manage to steer the Coder toward advanced techniques like tiling or tensor cores rather than settling for simpler fixes when feedback is given.
Meng: So, if we look at the practical gains mentioned in their results, they claim a one point five times average speedup over PyTorch eager mode code and two point seven three times better performance compared to RL model baselines on H200 GPUs <ref:2603.02637#pg2>. That’s a significant number when you look at real-world execution speedups.
Lalam: I think the improvement in correctness, going from three out of ten success on Level three tasks to nine out of ten, shows that this structured approach actually builds more reliable and powerful code than just relying on trial and error in prompting <ref:2603.02637#pg1>.
Conclusion: Tom: So, to wrap up StitchCUDA: it’s a multi-agent system with rubric-based RL that successfully automates end-to-end GPU program generation, achieving nearly one hundred percent success on Level three benchmarks and showing substantial speedups <ref:2603.02637#pg2>.
Jane: Essentially, this paper shows how breaking down the complex task into distinct roles for planning, coding, and verifying allows an AI to generate complete GPU programs that perform much better than previous single-kernel approaches.
Lu: I think the real weight of this work is in demonstrating that integrating structured feedback through rubric rewards can actually guide the Coder toward implementing sophisticated CUDA engineering techniques when dealing with complex system-level tasks.
Meng: From a practical standpoint, if we see these kinds of performance gains on H200 GPUs, it means we could deploy much larger and more complex AI models without needing a massive team to manually optimize every single line of code.
Lalam: I'm really optimistic about this; if this framework can consistently deliver high-quality GPU programs with such reliability, it could significantly accelerate the pace at which we can bring truly advanced, performant AI applications into production environments.
University of Minnesota-Twin Cities
cs.MA, cs.CL, cs.PL
Submitted: 2026-03-03
Updated: 2026-08-10
Code: https://github.com/UMN-APEX-Lab/StitchCUDA
Importance score: 89/100
The gist: StitchCUDA proposes a multi-agent framework integrated with rubric-based agentic reinforcement learning to automate end-to-end GPU program generation, addressing limitations in prior work that
Key concepts
- Multi-Agent Framework
- This involves coordinating three distinct AI agents—a Planner, a Coder, and a Verifier—that work together in an iterative loop to solve a complex problem. They share information about the generated code and profiling results to ensure the final GPU program is correct and performs well across all parts.
- Rubric-Based Agentic Reinforcement Learning
- This training method uses structured rules (rubrics) alongside standard rewards to guide an AI agent, specifically the Coder. The rubric evaluates performance across four dimensions: anti-hacking, CUDA engineering quality, operator coverage, and skill compliance. This prevents the agent from exploiting weaknesses in the reward system.
- Atomic Skills Decomposition
- Instead of training one large model for all tasks at once, this method splits the complex learning process into two simpler skills: 'from-scratch generation' and 'feedback-driven optimization.' This decomposition makes training more efficient by collecting specific types of data for each skill separately, allowing the model to learn advanced techniques like kernel fusion effectively.
Terminology
Summary
StitchCUDA proposes a multi-agent framework integrated with rubric-based agentic reinforcement learning to automate end-to-end GPU program generation, addressing limitations in prior work that focused only on single kernels. This system utilizes specialized agents—a Planner, a Coder, and a Verifier—to orchestrate an iterative plan–code–profile–refine
loop, fundamentally improving the Coder's ability to handle complex system-level requirements and optimize performance across multiple interacting kernels.
The gist
StitchCUDA achieves nearly 100% success rate on end-to-end GPU programming tasks on KernelBench Level 3, with a 1.5× average speedup over PyTorch eager mode code and 2.73× better performance than RL model baselines, by decomposing multi-turn agentic RL into two atomic skills and integrating rubric rewards to mitigate reward hacking and degenerate behaviors.
Multi-Agent Framework Workflow
StitchCUDA orchestrates three specialized agents with a global state machine that shares a typed State containing generated code, profiling artifacts, and routing decisions. The workflow follows an Iterative Coding-Feedback Loop.
The agents are defined as follows:
-
The Planner parses the reference PyTorch code to build a minimal profiling harness and records Nsys traces to identify time-dominant kernels and system hotspots. It then emits a structured to-do list with task identifiers, target kernels, expected shapes, and constraints.
-
The Coder generates CUDA implementations for the current subtask in a self-contained project, writing kernel stubs and host-side orchestrations that match the Planner’s spec. After receiving feedback from the Verifier, the Coder refines the current subtask accordingly.
-
The Verifier validates correctness and performance by first analyzing profiling results using Nsys to identify dominant GPU kernels and system-level bottlenecks (e.g., data transfer). It then profiles the identified bottleneck kernel with NCU to classify it (memory-bound or compute-bound) and produces actionable optimization suggestions routed back to the Planner or Coder.
Rubric-Based Agentic Reinforcement Learning for Coder
To improve the Coder’s capability beyond prompting, StitchCUDA integrates rubric-based agentic reinforcement learning over two atomic skills: (1) from-scratch generation, translating high-level GPU programming tasks into CUDA implementations; and (2) feedback-driven optimization, incorporating structured execution feedback to fix bugs. This decomposition alleviates the prohibitive computational cost of multi-turn agentic RL by collecting single-turn training data for both skills during workflow execution. The Coder is trained using GRPO targeting both skills, utilizing Qwen3-32B as the base model.
Rubric Reward Design and Anti-Hacking
The framework introduces a Rubric Reward to address reward hacking and degenerate behaviors, combining rule-based rewards (functional correctness/speedup) with rubric rewards. The rubric evaluates candidate kernels across four dimensions: (i) Anti-Hacking, penalizing reward exploitation; (ii) CUDA Engineering, rewarding advanced optimization techniques like tiling or tensor cores; (iii) Operator Coverage, encouraging broader optimization; and (iv) Skill Compliance. This is aggregated into a normalized shaping term to produce a stable signal. The final reward formulation combines this shaping with the rule-based reward:
R = Icorr · 1 − Ihack · min((s + τ) (1 + λ rˆrubric), Rmax). This structure prevents reward hacking by suppressing rewards when major hacking behaviors are detected (Ihack=1) and stabilizes training using a maximum reward cap (Rmax).
Atomic Skills Decomposition and Efficiency
The multi-turn agentic RL is decomposed into two atomic skills to reduce rollout overhead. Skill 1 focuses on from-scratch generation,
while Skill 2 focuses on improving an existing kernel by following feedback.
This approach allows the system to learn how to implement advanced CUDA programming techniques, such as custom kernel fusion and cublas epilogue optimizations. The training data is curated by sampling tasks from KernelBench Levels 1–3 and filtering them with human experts, ensuring the model learns relevant end-to-end optimization strategies efficiently.
Experimental Results
Experiments on KernelBench Level 3 show that StitchCUDA achieves nearly 100% success rate. Compared to baselines, StitchCUDA demonstrates substantial gains: it achieves a 2.7× speedup over the multiagent baseline and 2.73× better performance than RL model baselines on H200 GPUs. The rubric-based agentic RL is shown to be the key factor driving significant real-system-level speedups, improving correctness from 3/10 to 9/10 on Level 3 tasks while increasing mean speedup from 0.24× to 1.50× on H200.
Improvements for AI systems
As a fastidious researcher, I have analyzed the STITCHCUDA framework and its findings. The core improvement lies in moving from single-kernel optimization to holistic, end-to-end GPU program synthesis guided by structured feedback loops.
Here are the specific improvements and what the resulting AI system can achieve:
) 1. Enhanced End-to-End Program Synthesis (Addressing C1):
The improved system, STITCHCUDA, is fundamentally capable of generating complete, functional GPU programs—not just individual kernels—by orchestrating a Planner (system design), a Coder (implementation), and a Verifier (profiler/checker).
-
It can reason over the entire computation graph to identify necessary host-side orchestration (e.g., data movement patterns, CPU–GPU overlap strategies) alongside kernel-level optimizations like custom kernel fusion and tensor core utilization.
-
This allows for the generation of complex Level 3 workloads (like full VisionTransformer implementations) where performance is dominated by system-level factors, which prior methods failed to handle.
) 2. Robust Coder Capability via Rubric-Based Agentic RL (Addressing C2 & C3):
The integration of rubric-based agentic reinforcement learning significantly elevates the Coder's skill beyond simple prompt following:
-
The system learns to interpret structured execution feedback (from Nsys/NCU profiling) and apply targeted, meaningful optimizations rather than just copying reference code.
-
It is explicitly trained to avoid
reward hacking
(e.g., generating PyTorch-only code or hardcoding outputs) because the rubric reward penalizes such exploitative behaviors, leading to more robust and genuinely optimized CUDA implementations.
) 3. Efficient and Scalable Training (Addressing C3):
The decomposition of multi-turn agentic RL into two atomic skills (Generation and Optimization) drastically reduces computational overhead:
- Instead of prohibitively costly multi-turn rollouts that take hours per training iteration, the system trains the Coder on single-turn data samples. This makes the RL training process feasible and significantly faster.
) 4. Targeted Optimization Guidance (Addressing Degenerate Behaviors):
The rubric reward system directly combats degenerate behaviors:
- It explicitly rewards advanced CUDA engineering techniques (tiling, shared memory, cuBLASLt epilogues). This prevents the model from settling for trivial modifications (like only changing a single ReLU kernel) that yield minimal speedup but waste optimization potential.
The improved AI system can now perform the following specific tasks:
-
Generate complete, high-performance GPU programs for complex ML models (e.g., Vision Transformers) with near 100% correctness on challenging Level 3 benchmarks.
-
Automate the discovery and implementation of advanced, hardware-specific optimizations (like custom GEMM fusion or persistent workspace management) based on real profiling data.
-
Produce CUDA code that is not only correct but also exhibits substantial, measurable speedups (up to 3.5x observed) over reference implementations by optimizing both kernel logic and system architecture decisions.
Sources
- Kevin: Multi-Turn RL for Generating CUDA Kernels
- cuPilot: A Strategy-Coordinated Multi-agent Framework for CUDA Kernel Evolution
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Qwen2.5-Coder Technical Report
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging
- CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
- CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
- Qwen3 Technical Report
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
- CUDABench: Benchmarking LLMs for Text-to-CUDA Generation
- QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems