KernelFoundry: Hardware-aware evolutionary GPU kernel optimization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "KernelFoundry: Hardware-aware evolutionary GPU kernel optimization".
Jane: The paper was written by Nina Wiedemann, Quentin Leboutet, Michael Paulitsch, Diana Wofk and Benjamin Ummenhofer from Intel Corporation.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper discussion segment 1 — The Evolutionary Framework: Tom: Now, let's look at the core of the paper, specifically how KernelFoundry is summarized to achieve its goals. It’s not just iterative testing; they use a sophisticated methodology rooted in evolutionary algorithms.
Jane: The summary highlights this "Quality-Diversity" approach, which I think is a great concept to break down simply. Instead of just trying to find one single best kernel, the the system maintains a diverse archive of many high-performing candidates simultaneously.
Lu: That diversity is key because it prevents the entire population from collapsing onto one local optimum, which is a huge issue in any automated optimization search.
Meng: And this isn's not just theoretical; we can see how that translates to real-world use cases, like when optimizing for a specific model architecture such as Llama three point two 1B using their custom input format.
Lalam: I think the implication of this is that we are moving toward a computational infrastructure where the efficiency profile is actively managed and refined by machine learning principles.
Tom: Right, it’s not just about finding *a* solution, it’s about managing a vast landscape of possibilities to ensure we find the best one.
Jane: It's a powerful way to handle complexity by building that entire search space into the design itself.
Paper discussion segment 2 — Breakthrough Mechanisms: Tom: We've established that KernelFoundry is a sophisticated evolutionary framework, but what are the specific mechanisms that make it so effective? The authors introduce several breakthroughs beyond standard search techniques.
Jane: I’m really interested in the gradient-informed evolution. It suggests the system doesn't just randomly try things; it learns where to push next based on previous performance delta, which is a massive leap forward for directionality in optimization.
Lu: The calculation of those three gradients— grad F, grad R, and grad E—is what allows the powerful steering of the entire process. It tells the system not just *what* worked, but precisely how much to shift its focus based on empirical evidence.
Meng: From a practical standpoint, I’m impressed by their use of meta-prompt evolution to solve context degradation. That’s a massive practical solution for mitigating the LLM's tendency to get overwhelmed by its own failure history during iterative refinement.
Lalam: When Lalam thinks about this, it's about accelerating scientific discovery. If we can automate this level of optimization, researchers spend less time debugging hardware bottlenecks and more time exploring new ideas.
Tom: That’s the core idea—the automating the 'design' of a kernel from its 'tuning,' which is a massive gain in flexibility for a large engineering problem.
Jane: It’s taking specialized knowledge of GPU architecture and formalizing it into an automated, repeatable software structure that is something revolutionary.
Paper discussion segment 3 — Hardware-Awareness & Results: Tom: The next step is to talk about the results, specifically how KernelFoundry proves it is truly hardware-aware. It isn's just achieving a high average speedup; it understands the silicon.
Jane: The paper shows that when comparing its performance against baselines like "AI CUDA Engineer," the achieved speedup on tasks in the KernelBench subset is significantly higher, demonstrating genuine superiority in execution efficiency.
Lu: And this is tied to their use of a crossover experiment, which demonstrates that the kernel optimized for the target GPU outperforms kernels optimized on another hardware platform, confirming true hardware awareness.
Meng: From an engineering standpoint, I see this as a blueprint for building highly efficient AI services where you aren't limited by general-purpose benchmarks but are tailored to specific chip capabilities.
Lalam: The implication of this is that computational power is no longer limited by human optimization time, but by the ingenuity of our evolutionary tools.
Tom: It sounds like they've built an entire automated optimization engine rather than just providing a single algorithm, which really is the definition of a major advancement in this field.
Jane: It’s taking specialized knowledge and making it accessible through a robust framework, ensuring we are moving past "good enough" and into realizing highly specific efficiency gains.
Conclusion — Final Wrap-up: Tom: So, to wrap up our discussion of "KernelFoundry: Hardware-aware evolutionary GPU kernel optimization," it’s clear that achieving peak performance in AI computation requires a level of hardware understanding that moves far beyond traditional programming methods.
Jane: Exactly. The work fundamentally changes how we view the relationship between software design and silicon capabilities, creating a new standard for efficiency.
Lu: It’s truly exciting to see this systematic approach to optimizing computational bottlenecks across diverse hardware architectures, making the future of accelerated computing inevitable.
Meng: From a practical standpoint, this provides a clear blueprint for building deployable AI systems that are highly efficient and scalable in the real world.
Lalam: And that efficiency has massive downstream implications—it democratizes access to advanced computation globally by automating the tedious work of hardware tuning.
Tom: It really underscores that this is about making high-level AI capabilities accessible by optimizing the foundational plumbing underneath.
Jane: A perfect summary of the whole journey, I think.
Lu: This whole system is designed to change how we approach complex optimization problems through automated exploration and guided refinement.
Meng: It pushes the entire industry toward treating hardware optimization as a core, automated component of the AI pipeline itself.
Lalam: This structured approach ensures that we are not just achieving a single peak performance point, but discovering the entire landscape of optimal solutions simultaneously.
Intel Corporation
cs.DC, cs.LG
Submitted: 2026-03-12
Updated: 2026-09-03
Code: https://github.com/oneapi-src/oneDNN
Importance score: 89/100
The gist: KernelFoundry introduces a novel framework for "hardware-aware evolutionary GPU kernel optimization," addressing the critical challenge of maximizing computational efficiency in deep learning models.
Key concepts
- Quality-Diversity Approach
- This methodology maintains a diverse archive of many high-performing kernel candidates simultaneously. Instead of seeking one single best solution, it prevents the optimization process from getting stuck on a local optimum by exploring the entire search space.
- Gradient-Informed Evolution
- This mechanism allows the system to move beyond random testing. It learns the direction of improvement by calculating specific gradients ($ abla F, abla R, ext{and } abla E$) based on performance changes (delta). This enables precise steering of the optimization process.
- Hardware-Awareness
- The system is designed to understand specific silicon capabilities rather than just achieving high average speed. It confirms its superiority by optimizing kernels specifically for a target GPU, outperforming those that are optimized for different hardware platforms.
- Meta-prompt Evolution
- This is a practical solution used to mitigate the tendency of large language models (LLMs) to become overwhelmed by their own failure history during iterative refinement processes.
Terminology
Summary
KernelFoundry introduces a novel framework for hardware-aware evolutionary GPU kernel optimization,
addressing the critical challenge of maximizing computational efficiency in deep learning models. The paper details an advanced methodology that leverages evolutionary algorithms and hardware-specific knowledge to generate highly optimized GPU kernels, significantly improving runtime performance compared to existing state-of-the-art methods. This advancement is crucial for deploying large, complex AI models efficiently across diverse hardware architectures, thereby minimizing operational costs and expanding the feasibility of cutting-edge AI applications.
The Hardware-Aware Optimization Framework
KernelFoundry employs an evolutionary approach designed specifically to tailor kernel optimization to the target hardware. A key aspect of this framework is its ability to perform crossover-experiment[s],
where kernels are optimized in two separate runs on different GPUs—such as an Intel B580 and an Intel LNL GPU—and subsequently benchmarked on the respective other hardware. The authors emphasize that this cross-optimization capability yields superior results, noting that the kernel optimized for the target GPU outperforms the kernel that was optimized on the other GPU,
resulting in a hardware-speedup (hws) greater than 1.
Comparative Performance Analysis (KernelBench)
The framework's efficacy is demonstrated through rigorous benchmarking against established methods like OpenEvolve CodeLion [2025] using the KernelBench task set. The comparison across various operations, such as 16 ConvTranspose2d Mish Add Hardtanh Scaling and 99 Matmul GELU Softmax, reveals substantial performance gains. Analyzing the average runtimes shows that KernelFoundry achieves an average runtime of 3.325 ms on LNL, compared to 2.694 ms for OpenEvolve, and an average of 1.537 ms on B580, significantly outperforming the competitor's averages (1.036 ms). Furthermore, the calculated hardware-speedup (hws) metric provides quantitative proof of optimization superiority across all tested operations.
LLM Integration for Kernel Generation and Reproducibility
To demonstrate reproducibility and scalability, the authors utilized a large language model (LLM), GPT-OSS 20B, to generate optimized kernels. The process involved prompting the model to generate SYCL kernels for a representative subset of 20 kernels of KernelBench level 2.
While acknowledging that the model's lower capabilities led to failure in generating correct kernels in some cases, the successful optimizations achieved were highly impactful. For instance, using this methodology, the framework was able to achieve significant speedups for various operations:
-
16 ConvTranspose2d Mish Add Hardtanh Scaling.py: 1.775 Speedup -
21 Conv2d Add Scale Sigmoid GroupNorm.py: 2.106 Speedup -
35 Conv2d-Subtract-HardSwish-MaxPool-Mish.py: 1.576 Speedup
These results confirm the potential of integrating LLMs into the optimization pipeline, generating functional and high-performing kernels for complex operations like those found in modern deep learning architectures.
Improvements for AI systems
The core methodology presented—using evolutionary algorithms (EA) for hardware-aware GPU kernel optimization—is highly valuable. However, the current system relies heavily on manual selection of representative kernels and optimizing them in discrete runs. My proposed improvements focus on making the optimization process fully autonomous, generalizing hardware knowledge, and integrating real-time performance prediction into the model generation pipeline.
Improvement: Integrate a meta-learning layer that learns generalized kernel structures rather than optimizing specific named kernels (e.g., 16 ConvTranspose2d...). This system must move beyond the predefined list of 20 benchmark kernels and synthesize optimal implementations for any novel operation graph provided by the user or upstream model.
Mechanism:
-
Graph-to-Kernel Mapping: Instead of optimizing Operation X, the system optimizes the abstract computational graph G. The EA operates on parameterized representations of common graph nodes (e.g., a general
Conv2Dnode with adjustable group size, padding type, and activation function slot). -
Adversarial Hardware Probing: During training/optimization cycles, introduce an adversarial component that intentionally introduces minor variations (e.g., slightly different memory access patterns or data types) to the graph structure to force the EA to find robust, generalized optimizations rather than overfitting to specific test cases.
What the Improved System Can Do:
-
It can automatically and efficiently optimize custom model architectures (e.g., novel Mixture-of-Experts models, specialized graph neural networks) for any target GPU/accelerator without requiring manual selection of representative benchmark kernels.
-
It provides a
Kernel Optimization API
that accepts an arbitrary computational graph definition (e.g., ONNX or TorchScript format) and returns the optimized, compiled kernel code ready for deployment on the specified hardware, guaranteeing state-of-the-art performance relative to known bottlenecks.
Sources
- Kevin: Multi-Turn RL for Generating CUDA Kernels
- cuPilot: A Strategy-Coordinated Multi-agent Framework for CUDA Kernel Evolution
- CUDA-LLM: LLMs Can Write Efficient CUDA Kernels
- cuDNN: Efficient Primitives for Deep Learning
- STARK: Strategic Team of Agents for Refining Kernels
- AKG kernel Agent: A Multi-Agent Framework for Cross-Platform Kernel Synthesis
- NPUEval: Optimizing NPU Kernels with LLMs and Open Source Compilers
- Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
- PRAGMA: A Profiling-Reasoned Multi-Agent Framework for Automatic Kernel Optimization
- TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
- KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
- Illuminating search spaces by mapping elites
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- KernelBench: Can LLMs Write Efficient GPU Kernels?
- CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
- PEAK: A Performance Engineering AI-Assistant for GPU Kernels Powered by Natural Language Transformations
- SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
- Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
- Astra: A Multi-Agent System for GPU Kernel Performance Optimization
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing