LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching
AuthorsJinshan Liu, Haoran Qin, Xiaobing Tu, Jiacheng Liu, Jiahui Hu, Zhengan Yan, Yukun Xie, Kerui Shen, Jinkui Ren, Yuqi Lin, Xiantao Zhang, Linfeng Zhang
LinCa speeds up diffusion-based image and video generation by learning which parts of intermediate features can be safely reused, achieving substantial acceleration with minimal quality loss.
Key results
LinCa adds less than 0.2% parameters relative to the original diffusion model.
Maximum reported acceleration on Qwen-Image.
ImageReward achieved on Qwen-Image at 6.95× acceleration.
Reported acceleration on FLUX.1-dev.
VBench score on HunyuanVideo at 5.50× acceleration.
Additional speedup when LinCa is combined with INT8 quantization.
What the paper found
LinCa accelerates iterative diffusion inference by learning that hidden features evolve differently across models, denoising stages, and feature dimensions. Instead of applying one caching rule, it uses a lightweight fully invertible network to decompose cached features into three sub-components, reuses unstable components, and applies first- or second-order Hermite extrapolation to smoother components before losslessly reconstructing the original feature. Separate predictors are trained for timestep segments and target models using only pre-generated features, adding less than 0.2% parameters and requiring about 1 hour on a single 12GB GPU without loading diffusion-model weights. Across Alibaba Cloud’s Qwen-Image, FLUX.1-dev, and HunyuanVideo, LinCa reaches 6.95× acceleration on Qwen-Image with an ImageReward of 1.0524, 5.51× on FLUX.1-dev, and 5.50× on HunyuanVideo while preserving near-baseline quality. On HunyuanVideo, it records a VBench score of 80.16 versus 80.66 for the original model. The method also complements distillation and quantization, achieving 2.63× additional speedup on FLUX.1-dev-int8, and is evaluated on DrawBench, VBench, and GEdit-Bench. Unlike training-free approaches such as FORA, ToCa, and TaylorSeer, LinCa targets heterogeneous feature dynamics rather than imposing uniform prediction across the diffusion trajectory.
Original abstract
Diffusion models have achieved remarkable success in image and video generation, yet the high computational cost of iterative sampling remains a critical bottleneck for practical deployment. Feature caching has emerged as a promising acceleration paradigm by reusing or predicting intermediate features across timesteps. However, existing training-free methods apply uniform prediction strategies that cannot adapt to the heterogeneous feature dynamics, causing significant quality degradation under high acceleration ratios. We propose LinCa, a feature caching framework based on learnable invertible networks. LinCa decomposes cached features into sub-components with distinct continuity properties via a lightweight invertible network and applies differentiated prediction orders matched to each component. The strict invertibility guarantees lossless reconstruction back to the original feature space, forming a unified Decompose-Predict-Reconstruct pipeline. By training separate predictors for different models and timestep segments, LinCa adapts to heterogeneous feature dynamics. Experiments on FLUX, Qwen-Image, and HunyuanVideo demonstrate that LinCa, with less than 0.2% additional parameters, significantly outperforms existing methods and maintains near-lossless quality at 5-7x speedup. Code: https://github.com/QHR69/LinCa
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.