NTH

Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference

AuthorsZixuan Lan, Yanhong Li, Jiawei Zhou

August 25, 2026 2 min read
Watch on YouTube
The one-line take

RMM speeds up LLM inference by selectively skipping less informative parts of matrix multiplications while preserving much of the model's accuracy.

Key results

59.8
LLaMA 3.1 8B QA average at RR=0.5

Average zero-shot QA accuracy, compared with 69.8 for the full model.

37.5
CNN/DailyMail ROUGE-1 at RR=0.8

RMM summarization score, compared with 37.4 for the full LLaMA 3.1 8B model.

16.32
MLP Up accuracy drop at RR=0.7

Accuracy-point drop on ARC-Easy when reducing the MLP Up projection.

1.40
LLaMA 3.1 8B A100 speedup at 4096 tokens

End-to-end speedup from custom Triton kernels.

1.41
LLaMA 3.1 70B speedup at 2048 tokens

End-to-end speedup on an NVIDIA A100 at RR=0.8.

What the paper found

Reduced Matrix Multiplication, or RMM, is a training-free inference technique for Transformer models including Llama 3.1 and Qwen3 that dynamically selects the most informative dimensions of each matrix product using activation column norms and TopK, then computes only the retained contraction axes. Unlike SparseGPT, Wanda, or token and KV-cache pruning, it leaves model weights unchanged and applies to projections, attention scores QKᵀ, and value aggregation PV. The retention ratio provides a direct accuracy–compute control across models from 1B to 70B parameters. On LLaMA 3.1 8B at RR=0.5, average zero-shot QA accuracy is 59.8 versus 69.8 for the full model, outperforming static pruning baselines; on CNN/DailyMail at RR=0.8, ROUGE-1 is 37.5 versus 37.4 for the full model. Ablations show strong structural asymmetry: at RR=0.7, reducing the MLP Up projection causes a 16.32-point accuracy drop, while attention-side reduction causes only 3.52 points, so deployment should prune attention more aggressively than MLPs. Custom Triton kernels on an NVIDIA A100 deliver 1.40× end-to-end speedup at 4096 tokens for LLaMA 3.1 8B, and 1.41× at 2048 tokens for LLaMA 3.1 70B. RMM also transfers to Qwen2.5-VL and remains effective in long-context and multimodal inference, although gains depend on sequence length, hardware, and kernel integration.

Original abstract

Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix multiplications. We propose Reduced Matrix Multiplication (RMM), a training-free, input-adaptive inference method that reduces Transformer matrix products by selecting informative slices along their contraction dimensions, without modifying model weights. Under a simple retention-ratio control, RMM provides a smooth and predictable accuracy-efficiency trade-off. Across language models ranging from 1B to 70B parameters, we find that reduction tolerance depends on the model family, task, component, and retention ratio, although it often improves with model scale. Under moderate reduction, RMM remains robust across the evaluated discriminative, autoregressive generation, and long-context settings. We further show that the same principle extends to multimodal vision-language inference. Mechanistic ablations reveal a structural asymmetry within Transformers: attention-side computations are substantially more reducible than MLP components. Finally, wall-clock benchmarks with custom kernels on an NVIDIA A100 show that these computational savings can translate into practical runtime gains, especially at longer sequence lengths. Together, these results position RMM as a scalable direction for input-adaptive inference-time optimization.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis