PoEM: Predicting RL Outcomes from Existing Policies
AuthorsKimia Hamidieh, Giannis Daras, Antonio Torralba
AffiliationsMIT CSAIL
Resources
PoEM aims to mix the behaviors of already RL-trained models to predict how they would perform under a new reward, without running expensive RL again.
Key results
Number of GRPO or DPO reward adapters used in the language-model basis.
Effective rank of centered log-ratios for the 20-adapter Qwen3-0.6B basis.
Effective rank of parameter updates for the same basis.
Upper end of reward-gain recovery for covered held-out rewards.
Share of held-out rewards for which PoEM was closer to the target than the strongest single expert.
Average held-out image-reward gain recovered by composing the other adapters.
What the paper found
PoEM, or Product of Experts Mixing, predicts the result of reinforcement-learning post-training without running a new RL job. It composes existing adapters in log-policy space: for language models, token log-ratios against a shared reference are combined with weights recovered by ridge regression from a small calibration set scored by the new reward; for diffusion models, denoiser predictions are combined with a correction term. The key theoretical result is exact product-of-experts composition when the target reward is a linear combination of existing rewards, while experiments show the weaker and more useful condition that target behavior lies near the span of existing log-policies. On Qwen3-0.6B, a basis of 20 GRPO or DPO adapters had effective log-policy rank 6.3, versus 19.4 for parameter updates, revealing substantial behavioral redundancy. A coverage score predicts whether an unseen reward is reachable: on held-out rewards, covered cases recovered 0.55 to 0.69 of the directly trained RL policy’s reward gain, compared with 0.22 to 0.43 for uncovered cases. With ten diverse reward-model adapters, PoEM was closer to the directly trained policy than the strongest single expert on 90% of held-out rewards. On Stable Diffusion v1.4, composing 12 of 13 image adapters recovered 67% of the held-out adapter’s average reward gain. The method therefore offers an inference-time preview of expensive reward optimization, but cannot recover behavioral directions absent from its policy basis.
Original abstract
Foundation models are post-trained with reinforcement learning (RL) to maximize specific rewards, such as human alignment, correctness, or instruction following. This post-training process is computationally intensive, sometimes unstable, and has to be run from scratch every time the reward model changes or when we want to combine multiple rewards. We hence ask: given a new reward function, is it possible to predict the RL outcomes without actually running RL on it? We answer this in the affirmative by introducing PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards. First, we show that if the new reward function can be written as a linear combination of existing ones, then the new policy in log-space can be written as a linear combination of the existing log-policies. Surprisingly, even in cases where the rewards are not linearly connected, we observe that often log-policies from RL training span an approximately low-rank subspace across rewards. To our benefit, the weighting coefficients for this combination can be estimated using only the reward or basis policy outputs on the samples. We turn these observations into an algorithm that takes post-trained models and a new reward function, and approximates the target RL policy without actually running any additional RL training. We experimentally validate our approach across synthetic and real rewards, spanning both text and image modalities.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.