TTPO: Test-Time Policy Optimization

Published
Source
arXiv
Paper number
1030
Field
LLMs / NLP
arXiv ID
2608.27448

Key points

  • On competition-level hard problems, it found an asymmetry: even when pseudo-answers were wrong 85% of the time, 79% of the rollouts that disagreed with them were actually wrong.
  • Using this asymmetry, it developed TTPO, which applies distillation to agreeing rollouts and group RL penalties to disagreeing rollouts.
  • Without any labels, it matched or exceeded label-based learning (OPSD) on 5 competition-level benchmarks.
  • In a purely test-time learning setting, it raised Qwen3-1.7B accuracy from 38.0% to 45.2%.
  • It delivered gains of 25.2–36.4 percentage points even without thinking mode (long reasoning), and training on one benchmark generalized to others.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)