The Spectral Neuron replaces opaque nonlinearities with controllable eigenvalue-based functions, aiming to make expressive neural models more transparent and mathematically shapeable.
Key results
Rows used in the Criteo Display Advertising Challenge scaling experiment.
Rows used in the HIGGS scaling experiment.
Matrix dimension of a spectral neuron that fit bivariate targets of complexity 13 quite well.
Complexity level of the bivariate synthetic target fit by the 15-dimensional model.
What the paper found
The Spectral Neuron proposes a middle ground between transparent linear models and opaque neural networks. Instead of multiplying features by scalar weights, it forms an affine matrix pencil, A(x) = A₀ + ΣxᵢAᵢ, with learned real symmetric matrices, and uses one eigenvalue, λₖ(A(x)), as the prediction. This creates explicit, controllable nonlinearity: the smallest eigenvalue is concave, the largest is convex, positive-semidefinite coefficient matrices enforce monotonicity, and the spectral norm of each Aᵢ provides a global bound on that feature’s influence. Eigenvectors and eigenspaces also yield signed local influence measures and tighter instance-specific bounds, including at repeated eigenvalues through Clarke subdifferentials. Internal eigenvalues can universally approximate continuous functions on compact domains, while increasing matrix dimension increases expressive power without changing the model’s interpretable structure. Training uses standard eigensolvers, automatic differentiation in PyTorch, and Adam; the proposed initialization maintains eigenvalue gaps and avoids simultaneous diagonalization, which would reduce the model to piecewise-linear behavior. Experiments show scaling benefits on synthetic functions, the Criteo Display Advertising Challenge with 45,840,617 rows, and HIGGS with 11,000,000 rows; a 15-dimensional neuron fit bivariate targets of complexity 13 quite well, while monotonicity improved data efficiency on shape-constrained tasks. The approach is computationally expensive because dense eigendecomposition scales cubically with matrix dimension, so it targets applications where coefficient transparency and guaranteed shape constraints matter more than state-of-the-art accuracy. The implementation was assisted by OpenAI Codex.
Original abstract
As machine learned models increase in complexity and expressive power, features of simpler models, such as intrinsic coefficient transparency and control over the shape of the modeled function are lost. On the one edge of the spectrum we have simple linear models that possess coefficient transparency, but have a limited expressive power. On the other edge we have neural networks, that have expressive power that improves with scaling, but are mostly opaque. In this work we develop the \emph{spectral neuron} concept: a scalar model given by $f(x)=λ_k (A_0 + A_1 x + ... + A_n x_n)$, with learned real symmetric matrices $A_0, ..., A_n$. The input enters the model through an affine matrix function, but the prediction is obtained by reading one of its eigenvalues. Thus, the model is nonlinear, but the source of nonlinearity is still mathematically explicit. This gives us a useful middle ground: the model can become more expressive as the matrix dimension grows, while retaining coefficient transparency through the learned matrices. For example, extremal eigenvalues yield convex or concave functions, semidefinite constraints on the coefficient matrices impose monotonicity, and the associated eigenspaces characterize local feature influence. We study coefficient transparency, feature-influence bounds, and shape-control properties of this model family, and then test whether it can be learned and scaled in practice. We develop a systematic study of this model family, bringing together spectral results from several mathematical literatures to characterize its expressivity, coefficient transparency, feature influence, and shape-control properties. Code available at https://github.com/alexshtf/spectral_neuron_paper.
Read the original paperMore in Neural Networks
Browse all 22 papers →End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery
Akira Ito, Takayuki Miura, Yosuke Todo
A new query-efficient technique makes it possible to steal the parameters of small black-box neural networks using only their predicted labels.
Retrieving Individual Stems from Music Mixtures with Slot Embeddings
David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein
Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.
The Linear Representation Hypothesis Needs a Group Action
Louie Hong Yao, Yuhao Li, Shengchao Liu
This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.