NTH

Empirical Variational Autoencoder

AuthorsKaede Shiohara

AffiliationsThe University of Tokyo

October 11, 2026 2 min read
Watch on YouTube
The one-line take

EVA makes VAEs generate high-quality images and sounds faster by replacing their fixed Gaussian latent prior with a learned, self-predictive one.

Key results

7.68
ImageNet FID

EVA result at 256×256 resolution.

178.60
ImageNet Inception Score

EVA result at 256×256 resolution.

86.4M
EVA inference parameters

Main benchmark comparison, excluding tokenizer parameters.

24.76
Normalized AR-Diffusion inference time

24-block AR-Diffusion compared with EVA at 1.00.

0.64
VGGSound FID

EVA result, compared with 0.54 for AR-Diffusion.

8.68
LlamaGen-B ImageNet FID

EVA-B scores 7.68 in the same causal-model comparison.

What the paper found

Empirical Variational Autoencoder (EVA) revisits VAEs as generators for continuous sequences. Rather than forcing each latent token toward a fixed standard Gaussian, EVA learns a causal Gaussian prior from encoded training data, reducing the mismatch between prior and posterior. Its full-attention encoder and causal Transformer decoder use a learned latent predictor that adds only a linear layer; sampling proceeds ancestrally in latent space, where latent marginalization supports multimodal outputs without discrete quantization or iterative diffusion. On ImageNet at 256×256, EVA achieves FID 7.68 and Inception Score 178.60, with 86.4M inference parameters and normalized generation time 1.00, compared with 24.76 for 24-block AR-Diffusion. On VGGSound, EVA reaches FID 0.64, close to AR-Diffusion’s 0.54. In a separate causal ImageNet comparison, EVA-B records FID 7.68 versus LlamaGen-B’s 8.68, using 124.6M inference parameters compared with LlamaGen-B’s 153.4M. The results suggest that learning dependencies between latent positions can improve generation fidelity while retaining efficient sampling and VAE reconstruction capability.

Original abstract

We present Empirical Variational Autoencoder, a general generative framework for continuous-valued (i.e., non-vector-quantized) sequences. EVA is based on the evidence lower bound of the Variational Autoencoder (VAE) but learns autoregressive latent priors empirically from training data, which can be implemented only by an additional single linear layer on top of VAEs. By replacing the conventional standard-Gaussian constraint with the self-predicted priors, EVA significantly alleviates the latent distribution gap between prior and posterior which is typically observed in conventional VAEs, and leads to high-fidelity ancestral sampling for sequential data generation. Extensive experiments on image and sound synthesis demonstrate that EVA achieves competitive generation quality with autoregressive diffusion baselines despite its much faster inference time.

Read the original paper

More in Generative Models

Browse all 69 papers →
02Generative Model

MEND: RL For Flow Models via Proximal Velocity Matching

Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik

MEND makes flow-model reward tuning more efficient by moving samples only when the reward gain justifies the size of the change.

Read analysis