At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference
AuthorsBowen Wang, Chi Zhang, Diyou Shen, Renzo Andri, Navaneeth Kunhi Purayil, Luca Benini
Resources
Ventaglio is a compact vector-processor hardware extension that makes sparse Transformer inference substantially faster by accelerating the indexed memory operations that conventional architectures handle inefficiently.
Key results
Maximum speedup over optimized RVV baselines on moderate-sparsity tensor-contraction kernels.
Additional area introduced by Ventaglio in the Spatz cluster.
Maximum end-to-end speedup for DuoGPT-pruned LLaMA-3-8B prefill.
Maximum end-to-end speedup during autoregressive decoding.
Upper bound of the practical weight-and-activation sparsity range evaluated.
What the paper found
Researchers from ETH Zurich, the University of Bologna, and Huawei present Ventaglio, a runtime-configurable sparse execution unit and RISC-V Vector extension for Transformer inference. The design targets Gustavson’s dataflow, where sparse activation indices trigger computation and compressed weight metadata directs indexed accumulation. Unlike matrix-centric 2:4 mechanisms in NVIDIA Blackwell, AMD MI300, and Arm processors, Ventaglio fuses metadata decoding, gather-accumulate-scatter, and address post-increment operations beside the vector arithmetic unit, avoiding software index translation and L1-backed irregular memory traffic. Integrated into the open-source Spatz vector cluster and implemented in 12 nm FinFET, Ventaglio reaches a peak 7.4× kernel speedup over optimized RVV baselines while adding only 3.1% cluster area. A performance-calibrated GVSoC model evaluates a 4×4 multi-cluster system running the DuoGPT-pruned LLaMA-3-8B model in FP16 with practical 40–60% dual sparsity. End-to-end performance improves by up to 5.25× during prefill and 3.16× during autoregressive decoding. Prefill benefits from compute reduction in matrix-matrix projections, whereas decoding remains memory-bound, with weight sparsity reducing HBM traffic but fragmented DMA bursts limiting bandwidth utilization. The central result is that modest ISA and datapath changes can move sparse tensor contractions close to the hardware roofline across both matrix-matrix and matrix-vector workloads.
Original abstract
Fine-grained weight pruning and activation sparsification have emerged as effective approaches for reducing the compute and memory cost of inference for Transformer models. In the moderate-sparsity regime, Gustavson's dataflow provides a natural execution model for exploiting both activation and weight sparsity on vector processors through metadata-driven indexed accumulation. However, existing RVV architectures lack native support for this pattern, forcing kernels to rely on software index decoding and L1-backed indexed memory operations that keep sparse tensor contractions far below their roofline performance bound. We present Ventaglio, a runtime-configurable sparse execution unit coupled with RVV ISA extensions that drives sparse tensor contractions toward their roofline through indexed gather-accumulate-scatter support. Integrated into an open-source vector processing cluster and implemented in 12nm FinFET, Ventaglio accelerates sparse tensor contraction kernels by $6.9\text{--}7.4\times$ over optimized RVV baselines, with only $3.1\%$ area overhead for a cluster of tightly-L1 coupled vector processing elements. We build a performance-accurate instruction-level model of the Ventaglio extension, calibrate it against RTL implementation, and leverage it for scale-out performance analysis on a large $4\times4$ multi-cluster system. Using a DuoGPT-pruned LLaMA-3-8B model with practical $40\text{--}60\%$ dual sparsity, Ventaglio achieves $2.40\text{--}5.25\times$ and $2.06\text{--}3.16\times$ speedup over dense baselines during prefill and autoregressive decoding, respectively.
Read the original paperMore in AI Hardware
Browse all 34 papers →AI as a Compiler: Compiling Triton kernels without the Triton compiler
François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini
An LLM learns to replace parts of the GPU compiler stack by translating Triton code directly into fast, verified PTX kernels.
Purlin: Separating Orchestration from the Datapath of Collectives
Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis
Purlin makes GPU collective communication more modular and faster, improving large-scale LLM and diffusion inference across modern hardware.
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
Ashkan Vedadi Gargary, Guido Martínez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda, Madan Musuvathi, Tyler Sorensen
RESOLVE makes AI-written GPU kernels safer by combining race-finding tests with formal proofs that optimized code still computes the right result.