NTH

Auditing the Risk Claims of Distributional Reinforcement Learning

AuthorsHari Prasad

July 21, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that many supposedly meaningful 'risk-aware' signals learned by distributional RL agents are often just training artifacts, not real properties of the environment.

Key results

40–95%
Top-claim refutation range

Share of strongest risk claims refuted across QR-DQN, C51, and IQN on MinAtar.

84%
QR-DQN Breakout refutation

Strongest QR-DQN risk claims refuted on Breakout.

2,000
Audit rollout count

Snapshot-restart Monte Carlo futures sampled per action and audited state.

97%
RISKYGRID confirmations

Known trade-offs correctly confirmed by the positive-control audit.

0.92
RISKYGRID learned-to-true correlation

Correlation between learned and ground-truth risk magnitude.

34%
QR-DQN Seaquest CVaR safer-action rate

Share of flagged Seaquest states where the head’s CVaR choice selects the truly safer action.

What the paper found

Hari Prasad’s paper audits whether distributional reinforcement-learning heads actually represent the risk they appear to show. Using the excess Wasserstein gap between the top two actions, snapshot-restart Monte Carlo with 2,000 rollouts per action, permutation nulls, bootstrap refutation, and Benjamini–Hochberg FDR control, the study tests QR-DQN, C51, and IQN on MinAtar’s Breakout, Seaquest, and Asterix. Among the strongest learned risk claims, 40–95% are refuted, with QR-DQN specifically showing 84% refutation in Breakout and Seaquest and 66% in Asterix; essentially none of 245 audited states per game produces a statistically confirmed trade-off. The strongest claims are no better placed than a truth-blind ranking, and the artifact appears by 500k training steps, remains unrelated to final score, and varies by random seed. A positive-control environment, RISKYGRID, shows the audit can recover real risk: 97% of known trade-offs are confirmed, with learned-to-true correlation 0.92. Decision-making is consequential: QR-DQN’s CVaR choice selects the safer action in 81% of flagged Breakout states but only 34% in Seaquest, below the 63% achieved by mean-greedy selection. Risk-sensitive training, five-member ensembling, and recalibration do not restore reliable claims; recalibration succeeds on MinAtar only by shrinking the reported signal into near silence. The conclusion is that, for these agents, distributional heads often encode training artifacts rather than environment stochasticity, so risk-sensitive control and interpretability require ground-truth auditing before deployment.

Original abstract

Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but that has not been measured directly: are the risk claims of a trained distributional agent true? Our audit combines a decision-relevant screening metric (the excess Wasserstein gap between the top two actions, which equals the mass by which first-order stochastic dominance is violated), ground truth from snapshot-restart Monte Carlo, and a statistical harness (permutation nulls, bootstrap refutation, FDR control) without which the audit itself manufactures false conclusions. Across QR-DQN, C51, and IQN on MinAtar (33 runs), 40-95% of the strongest claimed risk trade-offs are refuted at 95% confidence, the placement of the strongest claims is statistically indistinguishable from truth-blind, and essentially no claim is confirmable: for these agents, the learned "risk" reflects a training artifact rather than environment stochasticity. The artifact is structural (fully formed early in training, uncorrelated with final score, idiosyncratic to each seed) and appears unchanged at full-Atari scale, with every top Breakout claim of a pretrained near-state-of-the-art QR-DQN refuted. Positive controls of known magnitude confirm 96-100% of real claims (correlation 0.89-0.92): the reading measures the agents, not the audit. Acting on the heads' CVaR advice at their most-flagged states ranges from beneficial to significantly worse than chance. Neither training for risk nor ensembling removes the artifact, and recalibration passes the audit only by nullifying the claims: the head is uninformative, not merely miscalibrated. We release the toolkit and document two silent pitfalls that produced convincing but wrong audits of our own.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →