NTH

Dataset Complexity Shapes Finite-Distance Loss Geometry in Neural Networks

AuthorsJaeyong Bae, Hawoong Jeong

September 1, 2026 2 min read
Watch on YouTube
The one-line take

The study shows that more complex datasets reshape where low-loss parameter neighborhoods contract around trained neural networks.

Key results

512
Synthetic dataset size

Points used in each controlled synthetic dataset.

22
Maximum neighborhood scale

Largest nearest-neighbor neighborhood used in CMS.

2545
Synthetic-network parameters

Trainable parameters in the synthetic-data network.

2461
MNIST-network parameters

Trainable parameters in the downsampled-MNIST network.

3.508
Random-label symmetric hardening

Maximum finite-distance response increase for randomized MNIST labels.

What the paper found

This paper shows that dataset structure, not merely dataset size or input statistics, reshapes the finite-distance geometry of neural-network loss landscapes. The researchers measure multiscale local label mixing with CMS and estimate Franz–Parisi-inspired local entropy around zero-training-error solutions using adaptive sequential Monte Carlo. In a controlled synthetic experiment with 512 points and neighborhoods extending to kmax = 22, increasing label interweaving causes a sharper, earlier contraction of the effective low-loss solution volume near the reference; farther away, the radial loss-entropy derivative becomes weak and nearly common across conditions. The networks use two tanh hidden layers and contain 2545 trainable parameters for synthetic data and 2461 for downsampled MNIST. On MNIST, label randomization amplifies the bottleneck: at radius r = 1, shell-weighted training accuracy falls to 0.598 for the most complex synthetic condition compared with about 0.93 for low complexity, while natural digit-pair tasks show a milder but consistent trend, spanning CMS values from approximately 0.013 to 0.277 and accuracies of 0.979 for 0/1 versus 0.893 for 4/9. A symmetric finite-distance check also finds hardening of 0.417 for clean labels versus 3.508 for randomized labels, ruling out residual first-order tilt as the sole explanation. The central result is that complexity changes where solution neighborhoods contract and can alter their radial shape, something a single Hessian-based flatness measure cannot capture.

Original abstract

Finite datasets can share the same size and low-order statistics while differing strongly in structural complexity. We connect this dataset complexity to loss-landscape geometry by pairing local label mixing across neighborhood scales with local entropy around trained neural-network solutions. Adapted from the Franz--Parisi construction in spin-glass theory, local entropy measures the effective volume of low-loss, solution-like parameter configurations at each distance from a reference. We estimate it in finite networks using adaptive sequential Monte Carlo. In a controlled synthetic sweep, greater dataset complexity produces a larger decrease in local entropy near the reference. Farther away, its radial derivative becomes weak and nearly common across conditions. Dataset complexity therefore changes where the effective solution volume contracts, rather than making it decrease uniformly faster. Experiments on real image data show the same qualitative trend, with label randomization further amplifying the effect. These results show that dataset structure shapes how low-loss neighborhoods are organized across finite distances from trained solutions.

Read the original paper

More in Neural Networks

Browse all 22 papers →
02Neural Network

Retrieving Individual Stems from Music Mixtures with Slot Embeddings

David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein

Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.

Read analysis
03Neural Network

The Linear Representation Hypothesis Needs a Group Action

Louie Hong Yao, Yuhao Li, Shengchao Liu

This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.

Read analysis