Auditing the Risk Claims of Distributional Reinforcement Learning
AuthorsHari Prasad
Resources
This paper shows that many supposedly meaningful 'risk-aware' signals learned by distributional RL agents are often just training artifacts, not real properties of the environment.
Key results
Share of strongest risk claims refuted across QR-DQN, C51, and IQN on MinAtar.
Strongest QR-DQN risk claims refuted on Breakout.
Snapshot-restart Monte Carlo futures sampled per action and audited state.
Known trade-offs correctly confirmed by the positive-control audit.
Correlation between learned and ground-truth risk magnitude.
Share of flagged Seaquest states where the head’s CVaR choice selects the truly safer action.
What the paper found
Hari Prasad’s paper audits whether distributional reinforcement-learning heads actually represent the risk they appear to show. Using the excess Wasserstein gap between the top two actions, snapshot-restart Monte Carlo with 2,000 rollouts per action, permutation nulls, bootstrap refutation, and Benjamini–Hochberg FDR control, the study tests QR-DQN, C51, and IQN on MinAtar’s Breakout, Seaquest, and Asterix. Among the strongest learned risk claims, 40–95% are refuted, with QR-DQN specifically showing 84% refutation in Breakout and Seaquest and 66% in Asterix; essentially none of 245 audited states per game produces a statistically confirmed trade-off. The strongest claims are no better placed than a truth-blind ranking, and the artifact appears by 500k training steps, remains unrelated to final score, and varies by random seed. A positive-control environment, RISKYGRID, shows the audit can recover real risk: 97% of known trade-offs are confirmed, with learned-to-true correlation 0.92. Decision-making is consequential: QR-DQN’s CVaR choice selects the safer action in 81% of flagged Breakout states but only 34% in Seaquest, below the 63% achieved by mean-greedy selection. Risk-sensitive training, five-member ensembling, and recalibration do not restore reliable claims; recalibration succeeds on MinAtar only by shrinking the reported signal into near silence. The conclusion is that, for these agents, distributional heads often encode training artifacts rather than environment stochasticity, so risk-sensitive control and interpretability require ground-truth auditing before deployment.
Original abstract
Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but that has not been measured directly: are the risk claims of a trained distributional agent true? Our audit combines a decision-relevant screening metric (the excess Wasserstein gap between the top two actions, which equals the mass by which first-order stochastic dominance is violated), ground truth from snapshot-restart Monte Carlo, and a statistical harness (permutation nulls, bootstrap refutation, FDR control) without which the audit itself manufactures false conclusions. Across QR-DQN, C51, and IQN on MinAtar (33 runs), 40-95% of the strongest claimed risk trade-offs are refuted at 95% confidence, the placement of the strongest claims is statistically indistinguishable from truth-blind, and essentially no claim is confirmable: for these agents, the learned "risk" reflects a training artifact rather than environment stochasticity. The artifact is structural (fully formed early in training, uncorrelated with final score, idiosyncratic to each seed) and appears unchanged at full-Atari scale, with every top Breakout claim of a pretrained near-state-of-the-art QR-DQN refuted. Positive controls of known magnitude confirm 96-100% of real claims (correlation 0.89-0.92): the reading measures the agents, not the audit. Acting on the heads' CVaR advice at their most-flagged states ranges from beneficial to significantly worse than chance. Neither training for risk nor ensembling removes the artifact, and recalibration passes the audit only by nullifying the claims: the head is uninformative, not merely miscalibrated. We release the toolkit and document two silent pitfalls that produced convincing but wrong audits of our own.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.