Learning Suffers More Than the Policy Class Under Partial Observability: A Closed-Form Analysis
AuthorsIdil Gözel
Resources
In partially observed environments, reinforcement learning may fail not because the policy cannot represent the answer, but because the critic learns the wrong lesson.
Key results
Relative cost gap between the best memoryless policy and the full-state optimal controller.
Excess cost of the actor-critic equilibrium over the best memoryless policy.
Aliased critic curvature at the class optimum, versus the true conditional curvature of 1.11.
Closed-form resting gain, compared with the class-optimal gain 1.649.
Excess cost over the memoryless optimum at λ=0.999.
Factor by which the equilibrium gain changes across five disturbances with identical autocovariance.
What the paper found
This paper separates two failure modes in partially observable reinforcement learning: the policy-class gap, caused by limited expressiveness, and the learning gap, caused by biased optimization. In a solvable partially observed linear-quadratic control problem, the best memoryless policy is only 10.4% more costly than the full-state LQG controller, but standard actor-critic learning settles at a policy 35% worse than that memoryless optimum. The mechanism is an aliased TD(0)/LSTD critic: using exact quadratic features over the observed variable, it interprets persistence generated by an unobserved state as excessive value curvature, reporting 16.7 instead of the true conditional curvature 1.11, a 15-fold inflation. The resulting expected policy-gradient equilibrium is gain 24.79 rather than the class-optimal gain 1.649. Closed-form analysis shows the equilibrium scales as k_eq h equals a function of exploration-to-disturbance power, while its high-gain deployment cost converges to the disturbance power, 0.300 in the default setting, regardless of the disturbance distribution. Across five non-Gaussian disturbances with identical autocovariance, the equilibrium gain still varies by a factor of 3.4, showing that location depends on higher-order statistics even when cost does not. The proposed remedy is not more policy memory but a longer bootstrap horizon: with GAE(λ), the excess cost falls from 35.0% at λ=0 to 0.1% at λ=0.999, and the PPO experiments show that a 32-frame stack does not prevent the failure. The practical conclusion is to match the value-estimation horizon to the environment’s hidden-state memory.
Original abstract
When a reinforcement learning agent cannot observe the full state, we usually blame its policies: it cannot see enough to represent a good one. We show that in a solvable case the bigger problem lies elsewhere. Even when a good policy is available and the agent's value function is expressive enough to describe it exactly, learning still ends up somewhere far worse. We study a partially observed linear-quadratic problem in which a standard actor-critic learner can be solved in closed form. At our default setting the best policy the agent can represent is already close to optimal, costing 10.4% more than the ideal controller that observes everything. Learning does not find it. The algorithm instead comes to rest at a policy that is 35% worse than the best one available to it, and we can say exactly where and why. The cause is a bias in what the critic learns rather than a limit on what the actor can express. Because the agent cannot attribute what it sees to the part of the state it cannot observe, the critic misreads that unexplained variation as sharp curvature in its own value estimates, and the actor follows that error away from the optimum. We derive closed-form expressions for the resulting policy, for its cost, and for the one design choice that removes the problem, which is how far the learner looks ahead before trusting its own value estimates. Deep reinforcement learning experiments follow these predictions closely. Notably, giving the agent memory of past observations does not help, while changing how far it looks ahead does.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.