Tao Lu's research proposes an efficient GPU inference method for large language models with moderate weight sparsity. The approach uses a three-layer matrix storage format combining sparse tensor cores and CUDA cores to accelerate sparse matrix multiplication. Evaluations demonstrate it is the first to outperform dense matrix multiplication on modern GPUs with high-bandwidth memory, achieving up to 1.64x kernel-level speedup over SpInfer and up to 1.41x end-to-end speedup over FlashLLM.
No score is assigned. Sources and their independence are shown in the citation chain below.