Liquid Gated Attention
AuthorsYiheng Jiang, Yuanbo Xu, Yongjian Yang
Resources
LGA is an efficient attention mechanism that models irregular, long-range time-series dynamics without sequential numerical solvers.
Key results
Maximum sequence length evaluated in steps.
Average accuracy across six long-term classification datasets.
Relative reduction on the HR regression dataset versus the strongest baseline.
Inference throughput in sequences per second for 4,096-step sequences.
What the paper found
Liquid Gated Attention, or LGA, is a solver-free temporal operator for irregularly sampled and long-horizon time series. It derives an input-dependent decay gate from liquid time-constant dynamics, approximates its integral with learnable endpoint interpolation, and stores history in a matrix-valued fast-weight associative memory. This lets LGA incorporate observed time intervals directly while retaining parallel computation: causal encoding uses a prefix scan, and non-causal encoding uses associative matrix multiplication, both with linear complexity in sequence length. A sequence-level normalization bounds cumulative decay and changes gradient attenuation from exponential to polynomial, improving long-horizon stability. The resulting LFormer backbone combines decoupled multi-head LGA with residual SwiGLU channel mixers and was evaluated across six tasks and sixteen datasets, including sequences of up to 17,984 steps, against models such as Transformer, Mamba, R-ODE, NCDE, and ContiFormer. LFormer achieved 71.3% average accuracy on six long-term classification datasets, reducing heart-rate regression MAE by 47.72% relative to the strongest baseline and reducing interpolation MSE on the irregular USH climate dataset by 83.33%. On a 4,096-step throughput test, LFormer reached 4652.90 sequences per second, while preserving linear memory scaling. Ablations show that normalization, input-aware gating, output gating, and head-specific interpolation are especially important for sparse, noisy, and extrapolative settings; however, fixing interpolation at 0.5, the classical trapezoidal rule, improved the Pendulum tracking result by 23.90% in MSE.
Original abstract
Real-world time series often exhibit irregular sampling and extended temporal horizons, requiring models to capture continuous-time dynamics across arbitrary intervals without prohibitive scaling costs. Discrete-time methods collapse variable time intervals into static positional steps; solver-dependent continuous-time models preserve temporal structure but rely on sequential integration, precluding parallelization; and solver-free approximations avoid this cost yet none couples observed time intervals with input-driven state modulation. We propose Liquid Gated Attention (LGA), a solver-free parallel temporal operator. By parameterizing an input-driven gating mechanism with observed time intervals, LGA introduces a continuous-time inductive bias and formulates hidden state evolution as a fast-weight associative memory, enabling parallel computation across the temporal dimension. Using matrix associativity in non-causal encoding and a prefix scan in causal encoding, LGA attains linear temporal complexity in sequence length in both modes. A sequence-level normalization bounds cumulative temporal decay for stable long-horizon optimization. Building on LGA, we instantiate LFormer, a modular backbone for continuous-time representation learning. Across six tasks and sixteen datasets spanning up to 17,984 steps, LFormer demonstrates long-range dependency modeling, fine-grained state tracking, and trajectory reconstruction from sparse and noisy observations, while delivering competitive performance against state-of-the-art discrete-time and continuous-time baselines with linear scaling efficiency.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.