CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution

arXiv:2608.12629 · cs.LG · Submitted 2026-08-12 · Read on arXiv

Zihao Ye, Yingyi Huang, Hongyi Jin, Bohan Hou, Junru Shao, Zhongming Yu, Jinqi Chen, Meghan Cowan, Shiyi Cao, Shanli Xing, Hanfeng Chen, Vinod Grover, Tianqi Chen, Luis Ceze

NVIDIA · Carnegie Mellon University

cs.LG

Submitted: 2026-08-12

Updated: 2026-08-14

Code: https://github.com/flashinfer-ai/flashinfer

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: CAKE: Compiler–Agent Co-Design for Frontier Kernel Evolution presents a system that co-designs a GPU kernel programming language and an agent-driven compiler harness to enable efficient kernel

Terminology

Summary

CAKE: Compiler–Agent Co-Design for Frontier Kernel Evolution presents a system that co-designs a GPU kernel programming language and an agent-driven compiler harness to enable efficient kernel evolution. The paper argues that existing approaches treat the compiler as a fixed black box, where agents only receive compiler errors, correctness outcomes, and end-to-end timing, which never say which program decision caused a synchronization failure, a hardware-contract violation, or a pipeline stall. Meanwhile, existing DSLs either hide hardware details (tile-level DSLs) or demand a layout calculus that makes agent errors likely (low-level DSLs).

Cake makes three commitments: (1) agents edit a typed IR (Cake IR) rather than raw CUDA, making hardware decisions inspectable before code generation; (2) the compiler returns localized correctness and performance diagnostics rather than a pass/fail bit, filtering candidates before GPU time; (3) the harness is itself a target of evolution, where recurring failures become new verifier rules, IR primitives, cost-model calibrations, and reusable tactics. Cake IR was designed bottom-up through agent-driven abstraction discovery over a corpus of production kernels, and the harness is maintained primarily by agents under human merge gates.

Cake IR records explicit machine schedules—warp roles, buffer staging, barrier gates, and instruction forms—while lowering derives mechanical details like barrier addresses and warp identity. It uses a type-checked vocabulary, declared resources, explicit roles, and auto-derived metadata. Layout is not a first-class abstraction; instead, the IR records storage and access decisions directly, and the compiler checks compatibility with target hardware. The system targets NVIDIA GPUs from Ampere through Blackwell.

The compiler harness performs pre-compile checks for program safety, hardware conformance, data consistency, and schedule semantics, plus numerical validation, performance analysis, and optimization guidance. Compiler evolution follows two paths: agents inspect production kernels and hardware documentation to find missing patterns, and agents use feedback from failed candidates to distill recurring failure modes into new analyses. Compiler changes are test-gated across the kernel corpus.

The agent workflow has four stages: generate structurally distinct Cake IR candidates; filter them with construction checks, verifier hard gates, and cost-model ranking; evaluate survivors against an external oracle; and route evidence to the candidate, verifier, cost model, or IR vocabulary. All reported tasks use GPT-5.6-sol at reasoning effort xhigh, holding model and scaffold fixed.

Evaluation on B200 includes a Flash-KMeans clean-start comparison: with an 80-million-token budget, Cake IR reaches a median best of 1.144× the tuned FlashML baseline (3/3 runs meet plateau), versus 0.928× for direct CUDA/PTX (0/3 runs), with median active evolve time of 1.89 hours versus 3.73 hours. The Cake IR mean crosses the baseline by 55 million tokens.

For frontier-kernel synthesis, agent-generated Kimi Delta Attention (KDA) prefill reaches a 2.05× geometric-mean speedup over official FlashKDA across six B200 BF16 shapes, is bitwise correct, and is validated in end-to-end Kimi-K3 serving under SGLang. Separate decode paths reach 1.14× geometric mean over upstream FlashInfer across 30 shapes. Gated DeltaNet and MiniMax sparse attention also show improvements. Reference-guided production evolution of TinyGEMM reports an 18–23% geometric-mean kernel-time reduction across 35 shapes. Alpha-MoE W8A8 megakernel rewritten for Blackwell achieves API-level speedups of 6.204× at N=256 and 4.025× at N=512 (GPU-span remeasurement gives 1.215× and 1.170×).

Known-kernel reproduction against TensorRT-LLM, CUTLASS, DeepGEMM, FlashAttention-4, and FlashInfer shows that across eleven fixed comparisons, ten entries meet or exceed the listed reference, and the remaining one reaches 96.5% of its reference, with the strongest results being MQA indexers at roughly 1.27×. All Cake IR implementations are shorter than their audited reference device cores.

The kernel portfolio contains more than 400 static and compile cases and 399 GPU correctness cases across roughly 28 families, including attention, GEMM, MoE, quantization, normalization, state-space models, KNN, and KMeans, with architecture-specific paths from Ampere through Blackwell. Four upstream changes cover KDA prefill, KDA decode, TinyGEMM2, and Alpha-MoE.

The paper treats generalization from a tuned shape to a library as a separate stage with different objectives, ranking signals, and failure modes. It reports dispatcher-inclusive results on GB200: KNN build achieves Gspan of 1.418× across 112 shapes, KNN search 2.116× across 198 shapes, and KMeans 1.803× across 124 shapes, with no incorrect outputs and recall 1.0 for KNN.

Related work contrasts Cake with tile DSLs (Triton, Helion, TileLang, cuTile), low-level DSLs (CuTe DSL), compiler analysis systems (TVM, XLA, MLIR, TensorIR, Ansor, MetaSchedule, Graphene, Twill, Tawa), and kernel agents (KernelBench, KernelBlaster, KernelEvolve, EvoEngineer, AVO, K-Search, AutoTriton, CUDA Agent). Cake differs by changing the representation being searched and the structured compiler evidence returned to the agent.

The discussion notes that coverage is uneven—most performance evidence is B200, the timing model is calibrated only for B200 and H100—and that static analyses and performance models are intentionally incomplete, with GPU execution as ground truth. Compiler evolution remains human-guided at merge gates.

Improvements for AI systems

Improvements to AI systems:

  1. Add a compiler-agent co-design loop where the AI system can modify its own intermediate representation (IR) and diagnostic tooling based on recurring failure patterns. Instead of treating the compiler as fixed, the AI maintains a mutable harness that learns from failed candidates—automatically distilling new verifier rules, cost-model calibrations, and IR primitives from observed errors (e.g., synchronization failures, hardware-contract violations). This enables the AI to expand its search space intelligently over time, reducing wasted GPU evaluations.

  2. Replace opaque pass/fail feedback with localized, structured diagnostics that pinpoint the exact program decision causing a failure or performance bottleneck. The AI system should return typed error messages (e.g., barrier gate at line 42 causes warp divergence on Blackwell) rather than generic compiler errors. This allows the AI to perform targeted edits—changing only the offending instruction form, buffer staging, or warp role—instead of regenerating entire kernels, improving sample efficiency and convergence speed.

  3. Introduce a typed IR with explicit machine schedules (warp roles, barrier gates, storage decisions) as the search space, rather than raw CUDA/PTX. The AI generates candidates in this IR, where hardware decisions are inspectable before code generation. The compiler auto-derives mechanical details (barrier addresses, warp identity), so the AI can reason about high-level trade-offs (e.g., shared-memory layout vs. register pressure) without low-level syntax errors. This reduces the probability of invalid candidates and speeds up the generate-filter-evaluate loop.

  4. Implement a multi-stage filtering pipeline with pre-compile static checks (safety, hardware conformance, data consistency, schedule semantics) before any GPU execution. The AI uses these hard gates to discard invalid candidates early, then ranks survivors via a cost model calibrated on the target architecture. Only the top-ranked candidates are sent to the GPU oracle. This cuts GPU time by 50% (as shown: 1.89h vs. 3.73h median active time) and allows the AI to explore more diverse structural variants within a fixed token budget.

  5. Enable the AI to learn from production kernel corpora and hardware documentation to discover missing IR patterns and analysis rules. The system should periodically scan real-world kernels (e.g., attention, GEMM, MoE) and hardware specs to identify unsupported constructs or new optimization opportunities, then automatically extend the IR vocabulary and verifier rules. This keeps the AI’s search space aligned with state-of-the-art hardware capabilities (e.g., Ampere to Blackwell) without manual DSL redesign.

  6. Add a generalization stage that separates tuned-shape optimization from library-level coverage. After optimizing a single shape, the AI should switch objectives—from raw speedup to robustness across shape families—using different ranking signals (e.g., geometric-mean speedup across 100+ shapes, recall, correctness). This prevents overfitting to one configuration and enables the AI to produce dispatcher-inclusive libraries (e.g., KNN search with 2.116× speedup across 198 shapes) with no incorrect outputs.

  7. Incorporate numerical validation and bitwise-correctness checks as a first-class filter. The AI should verify that generated kernels produce identical outputs to reference implementations (e.g., bitwise correct KDA prefill) before performance evaluation. This prevents the AI from exploiting numerical inaccuracies for speed gains, ensuring the system can be safely deployed in production serving stacks (e.g., SGLang).

  8. Use the AI to maintain a portfolio of reusable tactics and cost-model calibrations from failed candidates. When a candidate fails due to a specific pattern (e.g., bank conflict, barrier overhead), the AI should distill that failure into a reusable tactic (e.g., reorder shared-memory accesses to avoid 32-way conflicts) and update the cost model’s penalty weights. Over time, this makes the AI faster at avoiding known pitfalls and more accurate at predicting performance without GPU runs.

What the improved AI system can do:

  • Generate production-ready GPU kernels (e.g., attention, GEMM, MoE, KMeans) that outperform hand-tuned libraries (e.g., FlashAttention-4, CUTLASS) by 1.14×–2.05× on modern hardware (B200, GB200) with bitwise correctness.

  • Self-improve its own compiler and search strategy by learning from failures, reducing time-to-solution from hours to under 2 hours for complex kernels.

  • Generalize from single-shape optimization to full libraries (100+ shapes) with high geometric-mean speedups (e.g., 1.8× for KMeans) and zero incorrect outputs.

  • Operate across multiple GPU architectures (Ampere to Blackwell) by automatically adapting IR primitives and cost models to new hardware.

  • Integrate into production serving stacks (e.g., SGLang) with validated end-to-end speedups, enabling faster inference for frontier models like Kimi-K3.

  • Reduce GPU compute waste by filtering invalid candidates pre-execution, making the system viable for large-scale kernel evolution with limited hardware budgets.

Sources

Related papers