NTH

Neuronal Stochastic Attention Circuit (NSAC) for Probabilistic Representation Learning

AuthorsWaleed Razzaq, Yun-Bo Zhao

May 28, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a biologically inspired attention mechanism that turns attention scores into stochastic variables to produce uncertainty-aware predictions across several continuous-time tasks.

Key results

0.0002
Spiral MSE

On the irregular spiral function approximation benchmark, NSAC achieved the lowest MSE and outperformed Deep Ensembles and Monte-Carlo Dropout.

0.0095
Spiral CRPS

On the same spiral benchmark, NSAC achieved the best CRPS, indicating sharper probabilistic forecasts than the compared baselines.

-1.3816
ETTm1 NLL

On the ETTm1 long-range forecasting benchmark, NSAC achieved the best NLL among the evaluated methods.

0.0048
XJTU-SY MSE

On the XJTU-SY bearing prognostics benchmark, NSAC obtained the lowest MSE in the in-distribution setting.

0.0261
XJTU-SY CRPS

On XJTU-SY, NSAC also achieved the lowest CRPS, supporting strong probabilistic forecast quality.

What the paper found

The paper introduces the Neuronal Stochastic Attention Circuit, or NSAC, a continuous-time attention layer that replaces deterministic attention logits with the exact solution of an Ornstein–Uhlenbeck stochastic differential equation, using input-dependent gates borrowed from C. elegans Neuronal Circuit Policies. This design turns attention logits into a Gaussian process over time and pushes that uncertainty through a logistic-normal attention distribution, so the model produces probabilistic outputs natively rather than via post-hoc calibration. Training combines Gaussian negative log-likelihood with an epistemic-separation regularizer that increases variance on synthetic out-of-distribution perturbations, enabling explicit separation of aleatoric and epistemic uncertainty. Across five benchmarks—irregular spiral function approximation, Boston Housing and Kin8nm regression, ETTm1 and Jena-Climate forecasting, XJTU-SY/PRONOSTIA/HUST bearing prognostics, and Udacity plus OpenAI CarRacing control—NSAC is consistently competitive on accuracy and often stronger on uncertainty quality. On the spiral task it reaches MSE 0.0002 and CRPS 0.0095, outperforming Deep Ensembles and Monte-Carlo Dropout; on ETTm1 it achieves the best NLL at -1.3816 while maintaining reasonable forecast calibration; on XJTU-SY it obtains MSE 0.0048 and CRPS 0.0261. The key novelty is architectural: uncertainty is embedded inside continuous-time attention itself, with a closed-form forward pass that avoids numerical SDE solvers and remains interpretable at the neuronal cell level.

Original abstract

Reliable quantification of uncertainty estimates in continuous-time (CT) representation learning remains nascent, particularly within CT attention architectures. We introduce the Neuronal Stochastic Attention Circuit (NSAC), a novel biologically-inspired CT attention architecture that reformulates attention logit computation as the solution of an Ornstein-Uhlenbeck stochastic differential equation modulated by input-dependent, nonlinear interlinked gates derived from repurposed C.elegans Neuronal Circuit Policies (NCPs) wiring mechanism. It induces Gaussian distribution over logits that propagates principled stochasticity through logistic-normal distribution over attention weights to yield probabilistic output. A two-term objective function combining Gaussian negative log-likelihood with an epistemic-separation regularizer enforces higher predictive variance and enables joint quantification of aleatoric and epistemic uncertainty. Empirically, we implement NSAC in a diverse set of learning tasks including: (i) irregular CT function approximation; (ii) multivariate regression; (iii) long-range forecasting; (iv) Industry 4.0; and (v) the lane-keeping of autonomous vehicles. We observe that the NSAC remains competitive against several baselines in terms of accuracy and produces reasonably well-calibrated uncertainty estimates while being interpretable at the neuronal cell level.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis