Dataset Complexity Shapes Finite-Distance Loss Geometry in Neural Networks
AuthorsJaeyong Bae, Hawoong Jeong
Resources
The study shows that more complex datasets reshape where low-loss parameter neighborhoods contract around trained neural networks.
Key results
Points used in each controlled synthetic dataset.
Largest nearest-neighbor neighborhood used in CMS.
Trainable parameters in the synthetic-data network.
Trainable parameters in the downsampled-MNIST network.
Maximum finite-distance response increase for randomized MNIST labels.
What the paper found
This paper shows that dataset structure, not merely dataset size or input statistics, reshapes the finite-distance geometry of neural-network loss landscapes. The researchers measure multiscale local label mixing with CMS and estimate Franz–Parisi-inspired local entropy around zero-training-error solutions using adaptive sequential Monte Carlo. In a controlled synthetic experiment with 512 points and neighborhoods extending to kmax = 22, increasing label interweaving causes a sharper, earlier contraction of the effective low-loss solution volume near the reference; farther away, the radial loss-entropy derivative becomes weak and nearly common across conditions. The networks use two tanh hidden layers and contain 2545 trainable parameters for synthetic data and 2461 for downsampled MNIST. On MNIST, label randomization amplifies the bottleneck: at radius r = 1, shell-weighted training accuracy falls to 0.598 for the most complex synthetic condition compared with about 0.93 for low complexity, while natural digit-pair tasks show a milder but consistent trend, spanning CMS values from approximately 0.013 to 0.277 and accuracies of 0.979 for 0/1 versus 0.893 for 4/9. A symmetric finite-distance check also finds hardening of 0.417 for clean labels versus 3.508 for randomized labels, ruling out residual first-order tilt as the sole explanation. The central result is that complexity changes where solution neighborhoods contract and can alter their radial shape, something a single Hessian-based flatness measure cannot capture.
Original abstract
Finite datasets can share the same size and low-order statistics while differing strongly in structural complexity. We connect this dataset complexity to loss-landscape geometry by pairing local label mixing across neighborhood scales with local entropy around trained neural-network solutions. Adapted from the Franz--Parisi construction in spin-glass theory, local entropy measures the effective volume of low-loss, solution-like parameter configurations at each distance from a reference. We estimate it in finite networks using adaptive sequential Monte Carlo. In a controlled synthetic sweep, greater dataset complexity produces a larger decrease in local entropy near the reference. Farther away, its radial derivative becomes weak and nearly common across conditions. Dataset complexity therefore changes where the effective solution volume contracts, rather than making it decrease uniformly faster. Experiments on real image data show the same qualitative trend, with label randomization further amplifying the effect. These results show that dataset structure shapes how low-loss neighborhoods are organized across finite distances from trained solutions.
Read the original paperMore in Neural Networks
Browse all 22 papers →End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery
Akira Ito, Takayuki Miura, Yosuke Todo
A new query-efficient technique makes it possible to steal the parameters of small black-box neural networks using only their predicted labels.
Retrieving Individual Stems from Music Mixtures with Slot Embeddings
David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein
Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.
The Linear Representation Hypothesis Needs a Group Action
Louie Hong Yao, Yuhao Li, Shengchao Liu
This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.