PyramidFFT: Rearchitecting FFT with Matrix-Aligned Nested-Radix for Hierarchical Scratchpad Memory on AI Accelerators

Authors: Xiang Zhao, Ruge Zhang, Haipeng Jia, Kun Li, Jianliang Xu, Ting Cao, Yunxin Liu, Yunquan Zhang.

Published in: ACM Transactions on Architecture and Code Optimization (TACO), 2026

Abstract: Emerging AI accelerators offer high compute throughput and large scratchpad memories, but hierarchical scratchpad organization is poorly utilized by memory-bound FFT workloads. PyramidFFT is a memory-efficient FFT system co-designed for AI accelerators with large-capacity scratchpad memory. It combines a Memory-Aware Nested Network Design that reduces memory access frequency, Matrix-Aligned Butterfly Layout Optimization that adapts butterfly parallelism and data layout for Tensor Core Units, and Symmetry-Driven Real FFT Compression that removes conjugate-symmetric components. PyramidFFT achieves up to an 8.2x efficiency improvement over a baseline without memory-hierarchy mapping, bridging FFT memory behavior with the compute-dense design of modern AI accelerators.

BibTeX

@article{tacopyramidfft,
  title={PyramidFFT: Rearchitecting FFT with Matrix-Aligned Nested-Radix for Hierarchical Scratchpad Memory on AI Accelerators},
  author={Xiang Zhao and Ruge Zhang and Haipeng Jia and Kun Li and Jianliang Xu and Ting Cao and Yunxin Liu and Yunquan Zhang},
  journal={ACM Transactions on Architecture and Code Optimization (TACO)},

  year={2026}
}

Download Paper