TTPO: Test-Time Policy Optimization
- Published
- Source
- arXiv
- Paper number
- 1030
- Field
- LLMs / NLP
- arXiv ID
- 2608.27448
Key points
- On competition-level hard problems, it found an asymmetry: even when pseudo-answers were wrong 85% of the time, 79% of the rollouts that disagreed with them were actually wrong.
- Using this asymmetry, it developed TTPO, which applies distillation to agreeing rollouts and group RL penalties to disagreeing rollouts.
- Without any labels, it matched or exceeded label-based learning (OPSD) on 5 competition-level benchmarks.
- In a purely test-time learning setting, it raised Qwen3-1.7B accuracy from 38.0% to 45.2%.
- It delivered gains of 25.2–36.4 percentage points even without thinking mode (long reasoning), and training on one benchmark generalized to others.
Paper links
External research summaries. These are not HDATF publications or measured product results.