SPIKe: Generating Practical SpMV Kernels for Sparse Iterative Methods on GPUs

Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2026

Sparse iterative methods with static sparse matrices underpin many applications, with sparse matrix-vector multiplication (SpMV) as their primary performance bottleneck. Existing optimizations often require costly matrix-specific preprocessing whose overhead cannot be amortized as iteration counts decrease. We present SPIKe, a practical SpMV kernel generator for efficient sparse iterative methods across a wide range of iteration counts. SPIKe combines online thread-level parallelism adaptation, online warp-level asymmetric scheduling, and offline architecture-domain co-aware code generation. Across 2,294 matrices, SPIKe outperforms six baselines with negligible preprocessing overhead. In end-to-end evaluations, it achieves average speedups of 2.78x for PageRank and 4.26x for GMRES over GraphBLAS and PETSc.

Recommended citation: Luhan Wang, Haipeng Jia, Kun Li, Zongyuan He, Shengguo Li, Ting Cao, Yunxin Liu, Yunquan Zhang, Yifeng Chen. (2026). "SPIKe: Generating Practical SpMV Kernels for Sparse Iterative Methods on GPUs." SC.
Download Paper