NTH

LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching

AuthorsJinshan Liu, Haoran Qin, Xiaobing Tu, Jiacheng Liu, Jiahui Hu, Zhengan Yan, Yukun Xie, Kerui Shen, Jinkui Ren, Yuqi Lin, Xiantao Zhang, Linfeng Zhang

August 25, 2026 2 min read
Watch on YouTube
The one-line take

LinCa speeds up diffusion-based image and video generation by learning which parts of intermediate features can be safely reused, achieving substantial acceleration with minimal quality loss.

Key results

0.2%
Additional parameters

LinCa adds less than 0.2% parameters relative to the original diffusion model.

6.95×
Qwen-Image acceleration

Maximum reported acceleration on Qwen-Image.

1.0524
Qwen-Image ImageReward

ImageReward achieved on Qwen-Image at 6.95× acceleration.

5.51×
FLUX.1-dev acceleration

Reported acceleration on FLUX.1-dev.

80.16
HunyuanVideo VBench

VBench score on HunyuanVideo at 5.50× acceleration.

2.63×
FLUX.1-dev-int8 acceleration

Additional speedup when LinCa is combined with INT8 quantization.

What the paper found

LinCa accelerates iterative diffusion inference by learning that hidden features evolve differently across models, denoising stages, and feature dimensions. Instead of applying one caching rule, it uses a lightweight fully invertible network to decompose cached features into three sub-components, reuses unstable components, and applies first- or second-order Hermite extrapolation to smoother components before losslessly reconstructing the original feature. Separate predictors are trained for timestep segments and target models using only pre-generated features, adding less than 0.2% parameters and requiring about 1 hour on a single 12GB GPU without loading diffusion-model weights. Across Alibaba Cloud’s Qwen-Image, FLUX.1-dev, and HunyuanVideo, LinCa reaches 6.95× acceleration on Qwen-Image with an ImageReward of 1.0524, 5.51× on FLUX.1-dev, and 5.50× on HunyuanVideo while preserving near-baseline quality. On HunyuanVideo, it records a VBench score of 80.16 versus 80.66 for the original model. The method also complements distillation and quantization, achieving 2.63× additional speedup on FLUX.1-dev-int8, and is evaluated on DrawBench, VBench, and GEdit-Bench. Unlike training-free approaches such as FORA, ToCa, and TaylorSeer, LinCa targets heterogeneous feature dynamics rather than imposing uniform prediction across the diffusion trajectory.

Original abstract

Diffusion models have achieved remarkable success in image and video generation, yet the high computational cost of iterative sampling remains a critical bottleneck for practical deployment. Feature caching has emerged as a promising acceleration paradigm by reusing or predicting intermediate features across timesteps. However, existing training-free methods apply uniform prediction strategies that cannot adapt to the heterogeneous feature dynamics, causing significant quality degradation under high acceleration ratios. We propose LinCa, a feature caching framework based on learnable invertible networks. LinCa decomposes cached features into sub-components with distinct continuity properties via a lightweight invertible network and applies differentiated prediction orders matched to each component. The strict invertibility guarantees lossless reconstruction back to the original feature space, forming a unified Decompose-Predict-Reconstruct pipeline. By training separate predictors for different models and timestep segments, LinCa adapts to heterogeneous feature dynamics. Experiments on FLUX, Qwen-Image, and HunyuanVideo demonstrate that LinCa, with less than 0.2% additional parameters, significantly outperforms existing methods and maintains near-lossless quality at 5-7x speedup. Code: https://github.com/QHR69/LinCa

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis