KernelFoundry: Hardware-aware evolutionary GPU kernel optimization

arXiv:2603.12440 · cs.DC, cs.LG · Submitted 2026-03-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "KernelFoundry: Hardware-aware evolutionary GPU kernel optimization".

Jane: The paper was written by Nina Wiedemann, Quentin Leboutet, Michael Paulitsch, Diana Wofk and Benjamin Ummenhofer from Intel Corporation.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 1 — The Evolutionary Framework: Tom: Now, let's look at the core of the paper, specifically how KernelFoundry is summarized to achieve its goals. It’s not just iterative testing; they use a sophisticated methodology rooted in evolutionary algorithms.

Jane: The summary highlights this "Quality-Diversity" approach, which I think is a great concept to break down simply. Instead of just trying to find one single best kernel, the the system maintains a diverse archive of many high-performing candidates simultaneously.

Lu: That diversity is key because it prevents the entire population from collapsing onto one local optimum, which is a huge issue in any automated optimization search.

Meng: And this isn's not just theoretical; we can see how that translates to real-world use cases, like when optimizing for a specific model architecture such as Llama three point two 1B using their custom input format.

Lalam: I think the implication of this is that we are moving toward a computational infrastructure where the efficiency profile is actively managed and refined by machine learning principles.

Tom: Right, it’s not just about finding *a* solution, it’s about managing a vast landscape of possibilities to ensure we find the best one.

Jane: It's a powerful way to handle complexity by building that entire search space into the design itself.

Paper discussion segment 2 — Breakthrough Mechanisms: Tom: We've established that KernelFoundry is a sophisticated evolutionary framework, but what are the specific mechanisms that make it so effective? The authors introduce several breakthroughs beyond standard search techniques.

Jane: I’m really interested in the gradient-informed evolution. It suggests the system doesn't just randomly try things; it learns where to push next based on previous performance delta, which is a massive leap forward for directionality in optimization.

Lu: The calculation of those three gradients— grad F, grad R, and grad E—is what allows the powerful steering of the entire process. It tells the system not just *what* worked, but precisely how much to shift its focus based on empirical evidence.

Meng: From a practical standpoint, I’m impressed by their use of meta-prompt evolution to solve context degradation. That’s a massive practical solution for mitigating the LLM's tendency to get overwhelmed by its own failure history during iterative refinement.

Lalam: When Lalam thinks about this, it's about accelerating scientific discovery. If we can automate this level of optimization, researchers spend less time debugging hardware bottlenecks and more time exploring new ideas.

Tom: That’s the core idea—the automating the 'design' of a kernel from its 'tuning,' which is a massive gain in flexibility for a large engineering problem.

Jane: It’s taking specialized knowledge of GPU architecture and formalizing it into an automated, repeatable software structure that is something revolutionary.

Paper discussion segment 3 — Hardware-Awareness & Results: Tom: The next step is to talk about the results, specifically how KernelFoundry proves it is truly hardware-aware. It isn's just achieving a high average speedup; it understands the silicon.

Jane: The paper shows that when comparing its performance against baselines like "AI CUDA Engineer," the achieved speedup on tasks in the KernelBench subset is significantly higher, demonstrating genuine superiority in execution efficiency.

Lu: And this is tied to their use of a crossover experiment, which demonstrates that the kernel optimized for the target GPU outperforms kernels optimized on another hardware platform, confirming true hardware awareness.

Meng: From an engineering standpoint, I see this as a blueprint for building highly efficient AI services where you aren't limited by general-purpose benchmarks but are tailored to specific chip capabilities.

Lalam: The implication of this is that computational power is no longer limited by human optimization time, but by the ingenuity of our evolutionary tools.

Tom: It sounds like they've built an entire automated optimization engine rather than just providing a single algorithm, which really is the definition of a major advancement in this field.

Jane: It’s taking specialized knowledge and making it accessible through a robust framework, ensuring we are moving past "good enough" and into realizing highly specific efficiency gains.

Conclusion — Final Wrap-up: Tom: So, to wrap up our discussion of "KernelFoundry: Hardware-aware evolutionary GPU kernel optimization," it’s clear that achieving peak performance in AI computation requires a level of hardware understanding that moves far beyond traditional programming methods.

Jane: Exactly. The work fundamentally changes how we view the relationship between software design and silicon capabilities, creating a new standard for efficiency.

Lu: It’s truly exciting to see this systematic approach to optimizing computational bottlenecks across diverse hardware architectures, making the future of accelerated computing inevitable.

Meng: From a practical standpoint, this provides a clear blueprint for building deployable AI systems that are highly efficient and scalable in the real world.

Lalam: And that efficiency has massive downstream implications—it democratizes access to advanced computation globally by automating the tedious work of hardware tuning.

Tom: It really underscores that this is about making high-level AI capabilities accessible by optimizing the foundational plumbing underneath.

Jane: A perfect summary of the whole journey, I think.

Lu: This whole system is designed to change how we approach complex optimization problems through automated exploration and guided refinement.

Meng: It pushes the entire industry toward treating hardware optimization as a core, automated component of the AI pipeline itself.

Lalam: This structured approach ensures that we are not just achieving a single peak performance point, but discovering the entire landscape of optimal solutions simultaneously.

Intel Corporation

cs.DC, cs.LG

Submitted: 2026-03-12

Updated: 2026-09-03

Code: https://github.com/oneapi-src/oneDNN

Importance score: 89/100

The gist: KernelFoundry introduces a novel framework for "hardware-aware evolutionary GPU kernel optimization," addressing the critical challenge of maximizing computational efficiency in deep learning models.

Key concepts

Quality-Diversity Approach
This methodology maintains a diverse archive of many high-performing kernel candidates simultaneously. Instead of seeking one single best solution, it prevents the optimization process from getting stuck on a local optimum by exploring the entire search space.
Gradient-Informed Evolution
This mechanism allows the system to move beyond random testing. It learns the direction of improvement by calculating specific gradients ($ abla F, abla R, ext{and } abla E$) based on performance changes (delta). This enables precise steering of the optimization process.
Hardware-Awareness
The system is designed to understand specific silicon capabilities rather than just achieving high average speed. It confirms its superiority by optimizing kernels specifically for a target GPU, outperforming those that are optimized for different hardware platforms.
Meta-prompt Evolution
This is a practical solution used to mitigate the tendency of large language models (LLMs) to become overwhelmed by their own failure history during iterative refinement processes.

Terminology

Summary

KernelFoundry introduces a novel framework for hardware-aware evolutionary GPU kernel optimization, addressing the critical challenge of maximizing computational efficiency in deep learning models. The paper details an advanced methodology that leverages evolutionary algorithms and hardware-specific knowledge to generate highly optimized GPU kernels, significantly improving runtime performance compared to existing state-of-the-art methods. This advancement is crucial for deploying large, complex AI models efficiently across diverse hardware architectures, thereby minimizing operational costs and expanding the feasibility of cutting-edge AI applications.

The Hardware-Aware Optimization Framework

KernelFoundry employs an evolutionary approach designed specifically to tailor kernel optimization to the target hardware. A key aspect of this framework is its ability to perform crossover-experiment[s], where kernels are optimized in two separate runs on different GPUs—such as an Intel B580 and an Intel LNL GPU—and subsequently benchmarked on the respective other hardware. The authors emphasize that this cross-optimization capability yields superior results, noting that the kernel optimized for the target GPU outperforms the kernel that was optimized on the other GPU, resulting in a hardware-speedup (hws) greater than 1.

Comparative Performance Analysis (KernelBench)

The framework's efficacy is demonstrated through rigorous benchmarking against established methods like OpenEvolve CodeLion [2025] using the KernelBench task set. The comparison across various operations, such as 16 ConvTranspose2d Mish Add Hardtanh Scaling and 99 Matmul GELU Softmax, reveals substantial performance gains. Analyzing the average runtimes shows that KernelFoundry achieves an average runtime of 3.325 ms on LNL, compared to 2.694 ms for OpenEvolve, and an average of 1.537 ms on B580, significantly outperforming the competitor's averages (1.036 ms). Furthermore, the calculated hardware-speedup (hws) metric provides quantitative proof of optimization superiority across all tested operations.

LLM Integration for Kernel Generation and Reproducibility

To demonstrate reproducibility and scalability, the authors utilized a large language model (LLM), GPT-OSS 20B, to generate optimized kernels. The process involved prompting the model to generate SYCL kernels for a representative subset of 20 kernels of KernelBench level 2. While acknowledging that the model's lower capabilities led to failure in generating correct kernels in some cases, the successful optimizations achieved were highly impactful. For instance, using this methodology, the framework was able to achieve significant speedups for various operations:

  • 16 ConvTranspose2d Mish Add Hardtanh Scaling.py: 1.775 Speedup

  • 21 Conv2d Add Scale Sigmoid GroupNorm.py: 2.106 Speedup

  • 35 Conv2d-Subtract-HardSwish-MaxPool-Mish.py: 1.576 Speedup

These results confirm the potential of integrating LLMs into the optimization pipeline, generating functional and high-performing kernels for complex operations like those found in modern deep learning architectures.

Improvements for AI systems

The core methodology presented—using evolutionary algorithms (EA) for hardware-aware GPU kernel optimization—is highly valuable. However, the current system relies heavily on manual selection of representative kernels and optimizing them in discrete runs. My proposed improvements focus on making the optimization process fully autonomous, generalizing hardware knowledge, and integrating real-time performance prediction into the model generation pipeline.


Improvement: Integrate a meta-learning layer that learns generalized kernel structures rather than optimizing specific named kernels (e.g., 16 ConvTranspose2d...). This system must move beyond the predefined list of 20 benchmark kernels and synthesize optimal implementations for any novel operation graph provided by the user or upstream model.

Mechanism:

  • Graph-to-Kernel Mapping: Instead of optimizing Operation X, the system optimizes the abstract computational graph G. The EA operates on parameterized representations of common graph nodes (e.g., a general Conv2D node with adjustable group size, padding type, and activation function slot).

  • Adversarial Hardware Probing: During training/optimization cycles, introduce an adversarial component that intentionally introduces minor variations (e.g., slightly different memory access patterns or data types) to the graph structure to force the EA to find robust, generalized optimizations rather than overfitting to specific test cases.

What the Improved System Can Do:

  • It can automatically and efficiently optimize custom model architectures (e.g., novel Mixture-of-Experts models, specialized graph neural networks) for any target GPU/accelerator without requiring manual selection of representative benchmark kernels.

  • It provides a Kernel Optimization API that accepts an arbitrary computational graph definition (e.g., ONNX or TorchScript format) and returns the optimized, compiled kernel code ready for deployment on the specified hardware, guaranteeing state-of-the-art performance relative to known bottlenecks.

Sources

Related papers