HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization

Published
Source
arXiv
Paper number
986
Field
Distributed Systems
arXiv ID
2608.21157

Key points

  • It introduced hierarchical planning that first selects an implementation space among PyTorch operators, CUDA libraries, and handwritten CUDA.
  • Within the chosen space, it iteratively improves implementations using task contracts, profiling feedback, and expert knowledge, finding executable code even with a small candidate budget.
  • It ranked best or tied for best in 22 of 27 comparisons across three base models and three difficulty levels, and achieved a valid-implementation rate of 71.6% with a single candidate.
  • In a scientific-computing stencil case, it found a kernel 1.53 times faster than cuDNN after five iterations, showing potential beyond standard training operations.
  • Evaluation focused on NVIDIA A100 and limited precision settings, and stencil validation covered only one operator and one configuration, so generalization to other devices and scientific operations remains unverified.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)