NTH

Preisach Attention: A Hysteretic Model of Sequential Memory

AuthorsPiotr Frydrych

May 27, 2026 2 min read
Watch on YouTube
The one-line take

This paper replaces softmax attention with a hysteresis-based memory model that tracks extrema, aiming to give transformers a more efficient way to remember long histories.

Key results

O(1)
Turing completeness depth

A single-layer PAL-Transformer is Turing-complete with O(1) depth, simulating a two-stack pushdown automaton.

O(log n)
Hard-attention transformer depth

The cited hard-attention transformer framework requires O(log n) depth for Turing completeness, providing the comparison point for PAL.

O(n log n d)
PAL inference time

The paper reports total PAL inference cost as O(n log n d), versus O(n2 d) for standard attention.

O(kd)
PAL memory

PAL uses O(kd) memory, where k is the extremum stack depth, compared with O(nd) for standard attention.

H = 1
Vector PAL Turing completeness heads

The vector PAL variant (vPAL) is Turing-complete with a single head, using a two-dimensional signal projection.

What the paper found

Preisach Attention A Hysteretic Model of Sequential Memory introduces the Preisach Attention Layer (PAL), which replaces softmax attention with binary Preisach relays parameterized by learned activation/deactivation thresholds (α, β) and an extremum stack state that stores only alternating local maxima and minima. The core novelty is value-based, rate-independent memory: PAL depends on the sequence of extrema rather than token positions or temporal gaps, and the paper proves that this stack is a minimal sufficient statistic for all continuous causal rate-independent functionals. The authors show that a single-layer PAL-Transformer is Turing-complete with O(1) depth, simulating a two-stack pushdown automaton using a Cantor-depth encoding; in contrast, hard-attention transformers require O(log n) depth in the cited framework of Pérez et al. They also prove incomparability with standard transformers: PAL computes historical range, max-minus-min, in O(1) layers, while transformers can do exact random-access retrieval, fcopy(x, p)=xp, which PAL cannot because non-extremal positions are erased by wiping. A logical characterization is given via Extremum First-Order Logic (EFO), with PAL corresponding exactly to EFO and thus to a strict fragment of FO+Aggregate. Computationally, PAL inference is O(n log n d) time and O(kd) memory versus O(n2 d) and O(nd) for standard attention, and the paper proposes applications to long episodic memory, anomaly detection, and threshold-driven sequential decision-making.

Original abstract

We introduce the Preisach Attention Layer (PAL), a novel sequence modelling architecture grounded in the classical Preisach hysteresis operator from mathematical physics. PAL replaces the softmax attention mechanism with a binary relay operator parameterised by learned activation and deactivation thresholds, maintaining a stack of local extrema as its internal state. A single-layer PAL-Transformer with O(1) depth is Turing-complete under arbitrary precision arithmetic, achievable through simulation of a two-stack pushdown automaton -- in contrast to the O(log n) depth required by standard hard-attention transformers. Second, we prove that the function classes computable by PAL and by the transformer are incomparable: PAL computes historical range statistics in O(1) layers that require O(log n) layers for transformers, while transformers support random-access retrieval that PAL cannot perform without auxiliary state. The separating property is rate-independence -- PAL responds only to the sequence of local extrema, not to absolute token positions or temporal spacing. Third, we show that the extremum stack constitutes a minimal sufficient statistic of the input history for all rate-independent functionals, providing a formal analogue of the wiping property in classical hysteresis theory. PAL is thus an efficient architecture for tasks with long episodic memory and weak positional dependence, with O(n log n) total inference cost versus O(n^2) for standard attention.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis