TTPO: Test-Time Policy Optimization
AuthorsAozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu, Jun Xiao, Yueting Zhuang, Hua Yang, Qianglong Chen, Yongliang Shen
Resources
TTPO lets language models improve their reasoning at test time by learning from their own votes while selectively correcting likely mistakes.
Key results
Fraction of competition-level prompts where majority-vote pseudo-labels were incorrect.
Fraction of rollouts disagreeing with the pseudo-label that were nevertheless wrong.
Label-free average accuracy after TTPO, up from 38.0%.
Maximum average gain over the base model with thinking mode disabled.
Trajectories sampled for majority voting during TTPO training.
Optimal λ used to balance the GRPO and OPSD branches.
What the paper found
TTPO, or Test-Time Policy Optimization, enables label-free test-time training for mathematical reasoning by combining two asymmetric updates derived from majority-vote pseudo-labels. Instead of distilling every rollout toward a potentially wrong answer, it applies on-policy self-distillation, or OPSD, only to rollouts agreeing with the vote, while using Grouped Reinforcement Policy Optimization, or GRPO, to penalize disagreeing rollouts. Token-level weighting emphasizes uncertain or teacher-disagreeing positions in the distillation branch, and token masking penalizes only confident, anomalous errors in the reinforcement-learning branch. On Qwen3-1.7B, pseudo-labels were wrong on 85% of prompts, yet 79% of disagreeing rollouts were also wrong, supporting this asymmetric routing. Across AIME 2025, AIME 2026, HMMT 2025, HMMT 2026, and BRUMO 2025, TTPO matched or exceeded label-supervised OPSD. In fully label-free test-time training, it raised Qwen3-1.7B average accuracy from 38.0% to 45.2%, outperforming TTRL and OPSD-TTT; with thinking mode disabled, gains ranged from 25.2% to 36.4% across model scales. Training uses 64 sampled trajectories per problem, selects 8 for updates, and balances the GRPO branch with λ=0.1, producing a self-evolving loop in which improving votes generate stronger supervision.
Original abstract
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. We observe that this failure mode is asymmetric: rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct. Building on this observation, we propose Test-Time Policy Optimization (TTPO), an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL. Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors. Both updates remain well-grounded even under frequent pseudo-label errors, and majority-vote routing yields tighter self-supervision as the model improves. Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks, raises Qwen3-1.7B from 38.0% to 45.2% in TTT, yields +25.2% to +36.4% without thinking, and shows strong cross-task generalization.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.