NTH

Liquid Gated Attention

AuthorsYiheng Jiang, Yuanbo Xu, Yongjian Yang

September 9, 2026 2 min read
Watch on YouTube
The one-line take

LGA is an efficient attention mechanism that models irregular, long-range time-series dynamics without sequential numerical solvers.

Key results

17984
Longest sequence

Maximum sequence length evaluated in steps.

71.3%
Long-term classification accuracy

Average accuracy across six long-term classification datasets.

47.72%
Heart-rate MAE reduction

Relative reduction on the HR regression dataset versus the strongest baseline.

4652.90
LFormer throughput

Inference throughput in sequences per second for 4,096-step sequences.

What the paper found

Liquid Gated Attention, or LGA, is a solver-free temporal operator for irregularly sampled and long-horizon time series. It derives an input-dependent decay gate from liquid time-constant dynamics, approximates its integral with learnable endpoint interpolation, and stores history in a matrix-valued fast-weight associative memory. This lets LGA incorporate observed time intervals directly while retaining parallel computation: causal encoding uses a prefix scan, and non-causal encoding uses associative matrix multiplication, both with linear complexity in sequence length. A sequence-level normalization bounds cumulative decay and changes gradient attenuation from exponential to polynomial, improving long-horizon stability. The resulting LFormer backbone combines decoupled multi-head LGA with residual SwiGLU channel mixers and was evaluated across six tasks and sixteen datasets, including sequences of up to 17,984 steps, against models such as Transformer, Mamba, R-ODE, NCDE, and ContiFormer. LFormer achieved 71.3% average accuracy on six long-term classification datasets, reducing heart-rate regression MAE by 47.72% relative to the strongest baseline and reducing interpolation MSE on the irregular USH climate dataset by 83.33%. On a 4,096-step throughput test, LFormer reached 4652.90 sequences per second, while preserving linear memory scaling. Ablations show that normalization, input-aware gating, output gating, and head-specific interpolation are especially important for sparse, noisy, and extrapolative settings; however, fixing interpolation at 0.5, the classical trapezoidal rule, improved the Pendulum tracking result by 23.90% in MSE.

Original abstract

Real-world time series often exhibit irregular sampling and extended temporal horizons, requiring models to capture continuous-time dynamics across arbitrary intervals without prohibitive scaling costs. Discrete-time methods collapse variable time intervals into static positional steps; solver-dependent continuous-time models preserve temporal structure but rely on sequential integration, precluding parallelization; and solver-free approximations avoid this cost yet none couples observed time intervals with input-driven state modulation. We propose Liquid Gated Attention (LGA), a solver-free parallel temporal operator. By parameterizing an input-driven gating mechanism with observed time intervals, LGA introduces a continuous-time inductive bias and formulates hidden state evolution as a fast-weight associative memory, enabling parallel computation across the temporal dimension. Using matrix associativity in non-causal encoding and a prefix scan in causal encoding, LGA attains linear temporal complexity in sequence length in both modes. A sequence-level normalization bounds cumulative temporal decay for stable long-horizon optimization. Building on LGA, we instantiate LFormer, a modular backbone for continuous-time representation learning. Across six tasks and sixteen datasets spanning up to 17,984 steps, LFormer demonstrates long-range dependency modeling, fine-grained state tracking, and trajectory reconstruction from sparse and noisy observations, while delivering competitive performance against state-of-the-art discrete-time and continuous-time baselines with linear scaling efficiency.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis