NTH

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

AuthorsLuca Viano, Antoine Moulin, Audrey Huang, Volkan Cevher, Philip Amortila, Dylan J. Foster

August 9, 2026 3 min read
Watch on YouTube
The one-line take

This work shows that interacting with an expert can let a less expressive learner succeed by representing the expert’s values rather than exactly copying its policy.

Key results

4
Gymnasium environments

OVI was evaluated against BC, DAgger, and SPOIL in four Gymnasium environments.

64
Expert hidden-layer width

The expert network used 64 neurons per hidden layer.

2
Learner width range

The smallest learner networks used width 2, where OVI showed its strongest relative advantage.

50
Experimental seeds

Each Gymnasium comparison used 50 random seeds.

1/4
Offline VI suboptimality lower bound

Value-induced offline methods can suffer at least 1/4 suboptimality even with unlimited expert data.

What the paper found

This paper studies why querying an expert along the learner’s own trajectories can help when the learner cannot represent the expert’s full policy. Its central result is that on-policy interaction reduces the representational requirement from policy realizability to QπE-realizability: the learner only needs to represent the expert’s action-value function, not its exact action distribution. The authors introduce OVI, a layer-wise saddle-point algorithm that queries expert actions on learner-generated states, uses softmax policy updates, and requires only a linear maximization oracle over the value-function class. For finite classes, its expert-query complexity depends on log|Q|; convex value classes achieve an O(ε^-2) rate, while general nonconvex classes achieve O(ε^-4). The theory also establishes necessity: with horizon H=2, an offline learner can require Ω(|X|/ε) expert trajectories even when a value class of size 2 realizes QπE, and value-induced offline methods can retain at least 1/4 suboptimality with unlimited data. Experiments in four Gymnasium environments compare OVI with Behavior Cloning, DAgger, and SPOIL; an expert network uses 64 neurons per hidden layer, while learner widths range from 2 to 64. Across 50 seeds and 10 expert trajectories or query rounds, OVI performs best as learner capacity shrinks, supporting the claim that interaction and value-based learning provide complementary representational benefits. For chain-of-thought learning, OVI further replaces intractable response-level optimization over token sequences with tractable token-level updates, giving an exponential computational improvement under value realizability and motivating on-policy distillation and process reward models for language-model training.

Original abstract

Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches such as Behavior Cloning (BC) are known to suffer from compounding errors and performance plateaus, particularly when the learner cannot perfectly represent the expert's policy (as is typical, e.g., in distillation). Two interventions are widely understood empirically to improve performance: querying the expert interactively along the learner's own trajectories, and using value function estimation en route to generating a policy rather than directly fitting the expert's full action distribution. We investigate the nature of these improvements and their potentially surprising interplay. Our main finding is that expert interaction relaxes the representational demands on the learner: one only needs a model capable of realizing the expert's value function, bypassing the (often stricter) requirement of realizing the expert's policy itself. Concretely, we introduce OVI, an interactive on-policy IL algorithm that is statistically efficient whenever the learner can represent the expert's value function and computationally efficient given access to a linear maximization oracle. We complement this with a negative result showing that interaction is necessary. Namely, without stronger assumptions beyond expert-value realizability alone, any offline IL algorithm must scale with the complexity of the expert policy class. Our findings bear out empirically. OVI outperforms offline policy-based (BC), interactive policy-based (DAgger), and offline value-based IL methods, with the largest gains when the learner network is substantially less expressive than the expert's.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →