TTRL: Test-Time Reinforcement Learning

Published
Source
arXiv
Paper number
064
Field
RL / Inference
arXiv ID
2504.16084

Key points

  • Applying reinforcement learning to reasoning tasks for large language models is bottlenecked by the high cost and scarcity of human-annotated ground-truth labels needed for reward estimation.
  • Existing RL methods for LLMs rely heavily on supervised reward signals, which makes it difficult to keep improving on newly encountered or unlabeled test data.
  • The lack of explicit reward signals at inference time limits an LLM's ability to adapt to new distributions and improve generalization without human intervention.
  • It introduces Test-Time Reinforcement Learning, or TTRL, a framework that lets an LLM update its parameters and improve generalization at test time using RL.
  • It generates reward signals without explicit ground-truth labels by repeatedly sampling model outputs and deriving a consensus output through majority vote.
  • It computes a rule-based binary reward from whether each sampled output matches the majority-vote pseudo-label, and uses it to optimize the model policy with an RL algorithm such as GRPO.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)