HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization
- Published
- Source
- arXiv
- Paper number
- 986
- Field
- Distributed Systems
- arXiv ID
- 2608.21157
Key points
- It introduced hierarchical planning that first selects an implementation space among PyTorch operators, CUDA libraries, and handwritten CUDA.
- Within the chosen space, it iteratively improves implementations using task contracts, profiling feedback, and expert knowledge, finding executable code even with a small candidate budget.
- It ranked best or tied for best in 22 of 27 comparisons across three base models and three difficulty levels, and achieved a valid-implementation rate of 71.6% with a single candidate.
- In a scientific-computing stencil case, it found a kernel 1.53 times faster than cuDNN after five iterations, showing potential beyond standard training operations.
- Evaluation focused on NVIDIA A100 and limited precision settings, and stencil validation covered only one operator and one configuration, so generalization to other devices and scientific operations remains unverified.
Paper links
External research summaries. These are not HDATF publications or measured product results.