NTH

TTPO: Test-Time Policy Optimization

AuthorsAozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu, Jun Xiao, Yueting Zhuang, Hua Yang, Qianglong Chen, Yongliang Shen

September 2, 2026 2 min read
Watch on YouTube
The one-line take

TTPO lets language models improve their reasoning at test time by learning from their own votes while selectively correcting likely mistakes.

Key results

85%
Wrong pseudo-label prompts

Fraction of competition-level prompts where majority-vote pseudo-labels were incorrect.

79%
Wrong disagreeing rollouts

Fraction of rollouts disagreeing with the pseudo-label that were nevertheless wrong.

45.2%
Qwen3-1.7B TTT accuracy

Label-free average accuracy after TTPO, up from 38.0%.

36.4%
Non-thinking gains

Maximum average gain over the base model with thinking mode disabled.

64
Rollouts per problem

Trajectories sampled for majority voting during TTPO training.

0.1
GRPO loss weight

Optimal λ used to balance the GRPO and OPSD branches.

What the paper found

TTPO, or Test-Time Policy Optimization, enables label-free test-time training for mathematical reasoning by combining two asymmetric updates derived from majority-vote pseudo-labels. Instead of distilling every rollout toward a potentially wrong answer, it applies on-policy self-distillation, or OPSD, only to rollouts agreeing with the vote, while using Grouped Reinforcement Policy Optimization, or GRPO, to penalize disagreeing rollouts. Token-level weighting emphasizes uncertain or teacher-disagreeing positions in the distillation branch, and token masking penalizes only confident, anomalous errors in the reinforcement-learning branch. On Qwen3-1.7B, pseudo-labels were wrong on 85% of prompts, yet 79% of disagreeing rollouts were also wrong, supporting this asymmetric routing. Across AIME 2025, AIME 2026, HMMT 2025, HMMT 2026, and BRUMO 2025, TTPO matched or exceeded label-supervised OPSD. In fully label-free test-time training, it raised Qwen3-1.7B average accuracy from 38.0% to 45.2%, outperforming TTRL and OPSD-TTT; with thinking mode disabled, gains ranged from 25.2% to 36.4% across model scales. Training uses 64 sampled trajectories per problem, selects 8 for updates, and balances the GRPO branch with λ=0.1, producing a self-evolving loop in which improving votes generate stronger supervision.

Original abstract

Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. We observe that this failure mode is asymmetric: rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is correct. Building on this observation, we propose Test-Time Policy Optimization (TTPO), an asymmetric objective that distills agreeing rollouts via OPSD and penalizes disagreeing rollouts with Grouped RL. Token-level selection further refines both branches: distillation down-weights already-converged positions, while RL penalizes only confident errors. Both updates remain well-grounded even under frequent pseudo-label errors, and majority-vote routing yields tighter self-supervision as the model improves. Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks, raises Qwen3-1.7B from 38.0% to 45.2% in TTT, yields +25.2% to +36.4% without thinking, and shows strong cross-task generalization.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis