NTH

Learning Suffers More Than the Policy Class Under Partial Observability: A Closed-Form Analysis

AuthorsIdil Gözel

August 18, 2026 3 min read
Watch on YouTube
The one-line take

In partially observed environments, reinforcement learning may fail not because the policy cannot represent the answer, but because the critic learns the wrong lesson.

Key results

10.4%
Policy-class gap

Relative cost gap between the best memoryless policy and the full-state optimal controller.

35%
Learning gap

Excess cost of the actor-critic equilibrium over the best memoryless policy.

16.7
Critic curvature inflation

Aliased critic curvature at the class optimum, versus the true conditional curvature of 1.11.

24.79
Expected-update equilibrium gain

Closed-form resting gain, compared with the class-optimal gain 1.649.

0.1%
GAE horizon effect

Excess cost over the memoryless optimum at λ=0.999.

3.4
Non-Gaussian gain variation

Factor by which the equilibrium gain changes across five disturbances with identical autocovariance.

What the paper found

This paper separates two failure modes in partially observable reinforcement learning: the policy-class gap, caused by limited expressiveness, and the learning gap, caused by biased optimization. In a solvable partially observed linear-quadratic control problem, the best memoryless policy is only 10.4% more costly than the full-state LQG controller, but standard actor-critic learning settles at a policy 35% worse than that memoryless optimum. The mechanism is an aliased TD(0)/LSTD critic: using exact quadratic features over the observed variable, it interprets persistence generated by an unobserved state as excessive value curvature, reporting 16.7 instead of the true conditional curvature 1.11, a 15-fold inflation. The resulting expected policy-gradient equilibrium is gain 24.79 rather than the class-optimal gain 1.649. Closed-form analysis shows the equilibrium scales as k_eq h equals a function of exploration-to-disturbance power, while its high-gain deployment cost converges to the disturbance power, 0.300 in the default setting, regardless of the disturbance distribution. Across five non-Gaussian disturbances with identical autocovariance, the equilibrium gain still varies by a factor of 3.4, showing that location depends on higher-order statistics even when cost does not. The proposed remedy is not more policy memory but a longer bootstrap horizon: with GAE(λ), the excess cost falls from 35.0% at λ=0 to 0.1% at λ=0.999, and the PPO experiments show that a 32-frame stack does not prevent the failure. The practical conclusion is to match the value-estimation horizon to the environment’s hidden-state memory.

Original abstract

When a reinforcement learning agent cannot observe the full state, we usually blame its policies: it cannot see enough to represent a good one. We show that in a solvable case the bigger problem lies elsewhere. Even when a good policy is available and the agent's value function is expressive enough to describe it exactly, learning still ends up somewhere far worse. We study a partially observed linear-quadratic problem in which a standard actor-critic learner can be solved in closed form. At our default setting the best policy the agent can represent is already close to optimal, costing 10.4% more than the ideal controller that observes everything. Learning does not find it. The algorithm instead comes to rest at a policy that is 35% worse than the best one available to it, and we can say exactly where and why. The cause is a bias in what the critic learns rather than a limit on what the actor can express. Because the agent cannot attribute what it sees to the part of the state it cannot observe, the critic misreads that unexplained variation as sharp curvature in its own value estimates, and the actor follows that error away from the optimum. We derive closed-form expressions for the resulting policy, for its cost, and for the one design choice that removes the problem, which is how far the learner looks ahead before trusting its own value estimates. Deep reinforcement learning experiments follow these predictions closely. Notably, giving the agent memory of past observations does not help, while changing how far it looks ahead does.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →