When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning
AuthorsLuca Viano, Antoine Moulin, Audrey Huang, Volkan Cevher, Philip Amortila, Dylan J. Foster
Resources
This work shows that interacting with an expert can let a less expressive learner succeed by representing the expert’s values rather than exactly copying its policy.
Key results
OVI was evaluated against BC, DAgger, and SPOIL in four Gymnasium environments.
The expert network used 64 neurons per hidden layer.
The smallest learner networks used width 2, where OVI showed its strongest relative advantage.
Each Gymnasium comparison used 50 random seeds.
Value-induced offline methods can suffer at least 1/4 suboptimality even with unlimited expert data.
What the paper found
This paper studies why querying an expert along the learner’s own trajectories can help when the learner cannot represent the expert’s full policy. Its central result is that on-policy interaction reduces the representational requirement from policy realizability to QπE-realizability: the learner only needs to represent the expert’s action-value function, not its exact action distribution. The authors introduce OVI, a layer-wise saddle-point algorithm that queries expert actions on learner-generated states, uses softmax policy updates, and requires only a linear maximization oracle over the value-function class. For finite classes, its expert-query complexity depends on log|Q|; convex value classes achieve an O(ε^-2) rate, while general nonconvex classes achieve O(ε^-4). The theory also establishes necessity: with horizon H=2, an offline learner can require Ω(|X|/ε) expert trajectories even when a value class of size 2 realizes QπE, and value-induced offline methods can retain at least 1/4 suboptimality with unlimited data. Experiments in four Gymnasium environments compare OVI with Behavior Cloning, DAgger, and SPOIL; an expert network uses 64 neurons per hidden layer, while learner widths range from 2 to 64. Across 50 seeds and 10 expert trajectories or query rounds, OVI performs best as learner capacity shrinks, supporting the claim that interaction and value-based learning provide complementary representational benefits. For chain-of-thought learning, OVI further replaces intractable response-level optimization over token sequences with tractable token-level updates, giving an exponential computational improvement under value realizability and motivating on-policy distillation and process reward models for language-model training.
Original abstract
Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches such as Behavior Cloning (BC) are known to suffer from compounding errors and performance plateaus, particularly when the learner cannot perfectly represent the expert's policy (as is typical, e.g., in distillation). Two interventions are widely understood empirically to improve performance: querying the expert interactively along the learner's own trajectories, and using value function estimation en route to generating a policy rather than directly fitting the expert's full action distribution. We investigate the nature of these improvements and their potentially surprising interplay. Our main finding is that expert interaction relaxes the representational demands on the learner: one only needs a model capable of realizing the expert's value function, bypassing the (often stricter) requirement of realizing the expert's policy itself. Concretely, we introduce OVI, an interactive on-policy IL algorithm that is statistically efficient whenever the learner can represent the expert's value function and computationally efficient given access to a linear maximization oracle. We complement this with a negative result showing that interaction is necessary. Namely, without stronger assumptions beyond expert-value realizability alone, any offline IL algorithm must scale with the complexity of the expert policy class. Our findings bear out empirically. OVI outperforms offline policy-based (BC), interactive policy-based (DAgger), and offline value-based IL methods, with the largest gains when the learner network is substantially less expressive than the expert's.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.