NTH
AI research

Are Sparse Autoencoder Benchmarks Reliable?

AuthorsDavid Chanin

May 28, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that several popular sparse autoencoder benchmarks may be unreliable, and argues that the field needs better ways to tell good SAEs from bad ones.

Key results

16%–39%
TPP coefficient of variation

On the canonical Gemma Scope SAE, TPP has very high reseed noise at small top-N, making single-seed comparisons largely noise-dominated.

3%–7%
SCR coefficient of variation

On the canonical Gemma Scope SAE, SCR shows mid-single-digit reseed noise, with larger instability at bigger top-N settings.

ρ = −0.64
SCR top-500 synthetic correlation

On SynthSAEBench-16k, SCR at top-500 becomes negatively correlated with ground-truth MCC, meaning better SAEs can score worse on SCR.

0.2%
sae-probes reseed CV

The sae-probes sparse-probing variant is the most stable metric tested under reseed noise on the canonical SAE.

Spearman ρ = +0.87
sae-probes boolean in-sae synthetic correlation

On the synthetic boolean in-sae tasks, sae-probes at top-k = 16 aligns strongly with ground-truth MCC.

above 0.99
top-k = 16 accuracy

At canonical sparse-probing settings, top-k = 16 accuracy saturates on most trained SAEs even though GT-MCC still varies substantially.

What the paper found

David Chanin audits SAEBench, the standard evaluation suite for sparse autoencoders, and finds that benchmark reliability is substantially weaker than the field assumes. Using three lenses—five-reseed noise on canonical Gemma Scope, synthetic ground-truth correlation on SynthSAEBench-16k, and discriminability across 1.5B-token and 300M-token training trajectories—the paper shows that Targeted Probe Perturbation (TPP) and Spurious Correlation Removal (SCR) fail basic sanity checks at canonical settings. TPP often has 16%–39% coefficient of variation at small top-N and can worsen as training proceeds, while SCR reaches 3%–7% CV and becomes negatively correlated with ground truth at large top-N; on the synthetic panel, SCR top-500 flips to ρ = −0.64 versus ground-truth MCC. By contrast, sae-probes sparse probing is the most reliable metric tested, with reseed CV around 0.2% and strong synthetic calibration, reaching Spearman ρ = +0.87 against GT-MCC on boolean in-sae tasks. However, even sae-probes saturates quickly, with top-k = 16 accuracy above 0.99 on most trained SAEs despite GT-MCC still varying from 0.56 to 0.79. Across Matryoshka variants, most non-core metrics are dominated by seed and snapshot noise: only seven of 34 audited scores clear a 0.20 between-variant variance share, and canonical SCR top-10 and TPP top-10 sit near the boundary where a single-seed winner is often an accident of randomness. The central conclusion is blunt: SAEBench currently ranks SAEs reliably only when they are obviously different, so future progress needs larger task suites, more stable internal probes, and benchmark designs that reduce latent-sampling and probe-training noise.

Original abstract

Sparse autoencoders (SAEs) are a core interpretability tool for large language models, and progress on SAE architectures depends on benchmarks that reliably distinguish better SAEs from worse ones. We audit the SAE quality metrics in SAEBench, the de-facto standard SAE evaluation suite, through three complementary lenses: reseed noise on a fixed SAE, ground-truth correlation on synthetic SAEs, and discriminability across training trajectories. We find that two of these metrics, Targeted Probe Perturbation (TPP) and Spurious Correlation Removal (SCR), fail multiple lenses at their canonical settings and should not be used to evaluate SAEs. The other metrics show higher reseed noise and lower discriminability than the field assumes. The sae-probes variant of $k$-sparse probing is the most reliable metric we tested, but even sae-probes struggles to separate variants of the same SAE architecture. Our results show the field needs better SAE benchmarks.

Read the original paper