CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution

Published
Source
arXiv
Paper number
906
Field
Machine Learning
arXiv ID
2608.12629

Key points

  • When paired with a compiler that identifies what failed and why, the agent gets much closer to expert-level GPU kernels than it would in a black-box environment.
  • The harness itself is also evolved. Repeated failures are turned into verification rules, new IR instructions, and calibration of the performance model, so the system accumulates knowledge instead of using one-off workarounds.
  • Under the same budget, clean-start optimization with the Cake IR representation reaches 1.144 times the tuning baseline, while direct CUDA and PTX reach 0.928 times, which means the code representation itself matters.
  • The agent-generated Kimi Delta Attention kernel achieves a 2.05x geometric-mean speedup over the official implementation and passes real serving validation.
  • It completes library-level integration by separating single-target optimization from a roughly 400-target generalization, or dispatch, stage.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)