TTRL: Test-Time Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 064
- Field
- RL / Inference
- arXiv ID
- 2504.16084
Key points
- Applying reinforcement learning to reasoning tasks for large language models is bottlenecked by the high cost and scarcity of human-annotated ground-truth labels needed for reward estimation.
- Existing RL methods for LLMs rely heavily on supervised reward signals, which makes it difficult to keep improving on newly encountered or unlabeled test data.
- The lack of explicit reward signals at inference time limits an LLM's ability to adapt to new distributions and improve generalization without human intervention.
- It introduces Test-Time Reinforcement Learning, or TTRL, a framework that lets an LLM update its parameters and improve generalization at test time using RL.
- It generates reward signals without explicit ground-truth labels by repeatedly sampling model outputs and deriving a consensus output through majority vote.
- It computes a rule-based binary reward from whether each sampled output matches the majority-vote pseudo-label, and uses it to optimize the model policy with an RL algorithm such as GRPO.
Paper links
External research summaries. These are not HDATF publications or measured product results.