Preisach Attention: A Hysteretic Model of Sequential Memory
AuthorsPiotr Frydrych
Resources
This paper replaces softmax attention with a hysteresis-based memory model that tracks extrema, aiming to give transformers a more efficient way to remember long histories.
Key results
A single-layer PAL-Transformer is Turing-complete with O(1) depth, simulating a two-stack pushdown automaton.
The cited hard-attention transformer framework requires O(log n) depth for Turing completeness, providing the comparison point for PAL.
The paper reports total PAL inference cost as O(n log n d), versus O(n2 d) for standard attention.
PAL uses O(kd) memory, where k is the extremum stack depth, compared with O(nd) for standard attention.
The vector PAL variant (vPAL) is Turing-complete with a single head, using a two-dimensional signal projection.
What the paper found
Preisach Attention A Hysteretic Model of Sequential Memory introduces the Preisach Attention Layer (PAL), which replaces softmax attention with binary Preisach relays parameterized by learned activation/deactivation thresholds (α, β) and an extremum stack state that stores only alternating local maxima and minima. The core novelty is value-based, rate-independent memory: PAL depends on the sequence of extrema rather than token positions or temporal gaps, and the paper proves that this stack is a minimal sufficient statistic for all continuous causal rate-independent functionals. The authors show that a single-layer PAL-Transformer is Turing-complete with O(1) depth, simulating a two-stack pushdown automaton using a Cantor-depth encoding; in contrast, hard-attention transformers require O(log n) depth in the cited framework of Pérez et al. They also prove incomparability with standard transformers: PAL computes historical range, max-minus-min, in O(1) layers, while transformers can do exact random-access retrieval, fcopy(x, p)=xp, which PAL cannot because non-extremal positions are erased by wiping. A logical characterization is given via Extremum First-Order Logic (EFO), with PAL corresponding exactly to EFO and thus to a strict fragment of FO+Aggregate. Computationally, PAL inference is O(n log n d) time and O(kd) memory versus O(n2 d) and O(nd) for standard attention, and the paper proposes applications to long episodic memory, anomaly detection, and threshold-driven sequential decision-making.
Original abstract
We introduce the Preisach Attention Layer (PAL), a novel sequence modelling architecture grounded in the classical Preisach hysteresis operator from mathematical physics. PAL replaces the softmax attention mechanism with a binary relay operator parameterised by learned activation and deactivation thresholds, maintaining a stack of local extrema as its internal state. A single-layer PAL-Transformer with O(1) depth is Turing-complete under arbitrary precision arithmetic, achievable through simulation of a two-stack pushdown automaton -- in contrast to the O(log n) depth required by standard hard-attention transformers. Second, we prove that the function classes computable by PAL and by the transformer are incomparable: PAL computes historical range statistics in O(1) layers that require O(log n) layers for transformers, while transformers support random-access retrieval that PAL cannot perform without auxiliary state. The separating property is rate-independence -- PAL responds only to the sequence of local extrema, not to absolute token positions or temporal spacing. Third, we show that the extremum stack constitutes a minimal sufficient statistic of the input history for all rate-independent functionals, providing a formal analogue of the wiping property in classical hysteresis theory. PAL is thus an efficient architecture for tasks with long episodic memory and weak positional dependence, with O(n log n) total inference cost versus O(n^2) for standard attention.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.