Neuronal Stochastic Attention Circuit (NSAC) for Probabilistic Representation Learning
AuthorsWaleed Razzaq, Yun-Bo Zhao
Resources
This paper introduces a biologically inspired attention mechanism that turns attention scores into stochastic variables to produce uncertainty-aware predictions across several continuous-time tasks.
Key results
On the irregular spiral function approximation benchmark, NSAC achieved the lowest MSE and outperformed Deep Ensembles and Monte-Carlo Dropout.
On the same spiral benchmark, NSAC achieved the best CRPS, indicating sharper probabilistic forecasts than the compared baselines.
On the ETTm1 long-range forecasting benchmark, NSAC achieved the best NLL among the evaluated methods.
On the XJTU-SY bearing prognostics benchmark, NSAC obtained the lowest MSE in the in-distribution setting.
On XJTU-SY, NSAC also achieved the lowest CRPS, supporting strong probabilistic forecast quality.
What the paper found
The paper introduces the Neuronal Stochastic Attention Circuit, or NSAC, a continuous-time attention layer that replaces deterministic attention logits with the exact solution of an Ornstein–Uhlenbeck stochastic differential equation, using input-dependent gates borrowed from C. elegans Neuronal Circuit Policies. This design turns attention logits into a Gaussian process over time and pushes that uncertainty through a logistic-normal attention distribution, so the model produces probabilistic outputs natively rather than via post-hoc calibration. Training combines Gaussian negative log-likelihood with an epistemic-separation regularizer that increases variance on synthetic out-of-distribution perturbations, enabling explicit separation of aleatoric and epistemic uncertainty. Across five benchmarks—irregular spiral function approximation, Boston Housing and Kin8nm regression, ETTm1 and Jena-Climate forecasting, XJTU-SY/PRONOSTIA/HUST bearing prognostics, and Udacity plus OpenAI CarRacing control—NSAC is consistently competitive on accuracy and often stronger on uncertainty quality. On the spiral task it reaches MSE 0.0002 and CRPS 0.0095, outperforming Deep Ensembles and Monte-Carlo Dropout; on ETTm1 it achieves the best NLL at -1.3816 while maintaining reasonable forecast calibration; on XJTU-SY it obtains MSE 0.0048 and CRPS 0.0261. The key novelty is architectural: uncertainty is embedded inside continuous-time attention itself, with a closed-form forward pass that avoids numerical SDE solvers and remains interpretable at the neuronal cell level.
Original abstract
Reliable quantification of uncertainty estimates in continuous-time (CT) representation learning remains nascent, particularly within CT attention architectures. We introduce the Neuronal Stochastic Attention Circuit (NSAC), a novel biologically-inspired CT attention architecture that reformulates attention logit computation as the solution of an Ornstein-Uhlenbeck stochastic differential equation modulated by input-dependent, nonlinear interlinked gates derived from repurposed C.elegans Neuronal Circuit Policies (NCPs) wiring mechanism. It induces Gaussian distribution over logits that propagates principled stochasticity through logistic-normal distribution over attention weights to yield probabilistic output. A two-term objective function combining Gaussian negative log-likelihood with an epistemic-separation regularizer enforces higher predictive variance and enables joint quantification of aleatoric and epistemic uncertainty. Empirically, we implement NSAC in a diverse set of learning tasks including: (i) irregular CT function approximation; (ii) multivariate regression; (iii) long-range forecasting; (iv) Industry 4.0; and (v) the lane-keeping of autonomous vehicles. We observe that the NSAC remains competitive against several baselines in terms of accuracy and produces reasonably well-calibrated uncertainty estimates while being interpretable at the neuronal cell level.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.