Memorisation, convergence and generalisation in generative models
AuthorsAntoine Maillard, Sebastian Goldt
Resources
This paper shows that generative models can look converged without truly learning the underlying factors, separating surface-level matching from deeper representation recovery.
Key results
In the linear generative model, memorisation decays at large n and the model stops reproducing individual training points at small load.
The overlap between independently trained models becomes non-trivial when the number of samples is linear in the input dimension.
As sample complexity grows, the output overlap between two models trained on disjoint data converges to one.
The leading-eigenvector overlap for latent recovery shows a BBP transition and is zero below this sample-complexity threshold.
In the power-law image setting, the rotated overlap Q⋆ recovers a sharp phase transition near unit sample complexity.
What the paper found
Maillard and Goldt give an asymptotically exact theory of memorisation, convergence, and latent recovery in linear generative models, using a spiked Wishart Gaussian data model with covariance Σ⋆ = I + βuu⊤ and a regularised covariance estimator sampled through x̃(z)=Σ̂1/2 z. They show that memorisation, quantified by the maximum training-sample overlap m, is a small-load phenomenon: for large n it decays as m≈sqrt(2 log n/[n(1+σ²)]), so the model stops reproducing individual training points on the scale n=Θ(1). By contrast, the convergence overlap q between two independently trained models with the same latent z turns on only when n≍d, and is governed entirely by the Marchenko–Pastur bulk spectrum, with q→1 as γ=n/d→∞. The main novelty is that q does not detect whether the principal latent direction u has been recovered: that requires a separate overlap Q between top eigenvectors, which obeys a BBP transition at γc(β)=β−2 and is zero below threshold. They then connect q to the Kullback–Leibler divergence and Q to the maximum-sliced distance, showing that output agreement and latent-factor recovery are mathematically distinct objectives. Extending the analysis to power-law spectra, they explain why real images from CelebA and ImageNet make the naive Q misleading, and define a rotated overlap Q⋆ in Fourier space that recovers a sharp transition near γ≈1. In diffusion experiments with a U-Net on CelebA, q and Q⋆ reproduce the memorisation-to-generalisation transition and latent recovery behavior observed by Kadkhodaie et al. (ICLR 2024).
Original abstract
Generative neural networks learn how to produce highly realistic images from a large, but finite number of examples - or do they simply memorise their training set? To settle this question, Kadkhodaie, Guth, Simoncelli and Mallat (ICLR '24) trained diffusion models independently on disjoint subsets of a dataset and showed that they converge to nearly the same density when the number of training images is large enough. This result raises two basic questions: how much data do you need for convergence, and what does convergence capture about learning the data distribution? Here, we address these questions by providing an exact analytical characterisation of the transition from memorisation to generalisation in linear generative models. We find that these models memorise at small load, while convergence emerges continuously when the number of samples is linear in the input dimension. Strikingly, we find that convergence is insensitive to recovery of the principal latent factors of the data, which are recovered in a sharp transition. After extending our approach to data with power-law spectra, we find the same distinction between convergence and latent recovery in our experiments with convolutional denoisers and in the data of Kadkhodaie et al. We thus show that generalisation in generative models decomposes into at least two distinct objectives: matching the bulk of the data distribution and recovering the principal latent factors. These objectives correspond to two different distances between true and learnt data distribution, and only the first one is captured by convergence.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.