Prospective Coding Improves Learning in Deep Continuous-Time Recurrent Networks
AuthorsShivang Rawat, Mirko Morello, Flaviano Morone, David J. Heeger
Resources
This work shows that anticipating bottom-up signals can make deep continuous-time recurrent networks learn better with fewer parameters.
Key results
Six-layer width-32 RQF on raw-audio Google Speech Commands v0.02 under full BPTT.
Parameters used by the six-layer width-32 raw-audio RQF.
Prospective coding produced 10290 times larger first-layer weight gradients under spatial-only backpropagation.
Six-layer width-64 RQF on the 16,384-step Path-X benchmark.
Number of time steps in the Path-X input sequence.
What the paper found
This paper introduces Recursive Quadrature Filters, or RQFs: complex-valued, band-pass temporal filters that form a constrained diagonal state-space model and retain efficient zero-order-hold scans for long sequences. Its central intervention is prospective-input coding, which replaces each layer’s instantaneous bottom-up signal with a first-order look-ahead, χτ(x)=x+τẋ. The implementation is a parameter-free two-tap input path: it leaves the recurrent transition and parallel scan unchanged, adds only a lagged input read, and preserves asymptotic complexity. Theoretical analysis shows why this matters under spatial-only backpropagation: instantaneous inputs impose an explicit depth factor of (h/τ) per layer, while prospective inputs change the per-hop coefficient from O(h/τ) to O(1). Across RQF, S5, and nonlinear ORGaNICs, prospective variants match or improve controls under full BPTT and generally improve spatial-only learning. On raw-audio Google Speech Commands v0.02, a six-layer width-32 RQF reaches 96.09 ± 0.23% accuracy with 31.9k parameters; direct gradient measurements show 10290 times larger first-layer weight gradients than the instantaneous model. On the 16,384-step Path-X benchmark, a six-layer width-64 RQF reaches 83.56 ± 2.11%, exceeding chance by more than 30 points. The paper also distinguishes its signed two-tap correction from Mamba-3’s convex interpolation and demonstrates transfer to S5 and ORGaNICs.
Original abstract
Temporal integration gives continuous-time recurrent networks memory, but in deep stacks it also delays bottom-up signals and attenuates top-down errors. We develop Recursive Quadrature Filters (RQFs), biologically motivated complex-valued temporal filters that are a special case of diagonal state-space models (SSMs), and ask whether this failure mode can be addressed by making each layer's bottom-up input prospective. Starting from an energy model, we derive the RQF dynamics and show that each RQF is a band-pass filter whose learnable parameters control its tuning frequency and bandwidth. We then make each layer's bottom-up input prospective using a parameter-free two-tap update that leaves the recurrent transition and parallel scan unchanged. We extend this correction to general diagonal SSMs and show that it mitigates depth-dependent gradient attenuation when temporal gradients are truncated, i.e., spatial-only backpropagation. We evaluate the intervention in RQFs, S5, and ORGaNICs (a nonlinear gated RNN) trained using full backpropagation through time (BPTT) and spatial-only backpropagation. Under full BPTT, prospective variants match or outperform their non-prospective controls in every model and configuration. A non-residual width-32 six-layer RQF reaches 96.09% accuracy on raw-audio Speech Commands with 31.9k parameters; a width-64 six-layer RQF reaches 83.56% on the 16,384-step Path-X task. These results identify RQFs as a parameter-efficient recurrent substrate and prospective-input coding as an input-side correction for deep continuous-time recurrent networks.
Read the original paperMore in Neural Networks
Browse all 22 papers →End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery
Akira Ito, Takayuki Miura, Yosuke Todo
A new query-efficient technique makes it possible to steal the parameters of small black-box neural networks using only their predicted labels.
Retrieving Individual Stems from Music Mixtures with Slot Embeddings
David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein
Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.
The Linear Representation Hypothesis Needs a Group Action
Louie Hong Yao, Yuhao Li, Shengchao Liu
This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.