Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals

Published
Source
arXiv
Paper number
605
Field
Machine Learning
arXiv ID
2607.11505

Key points

  • By separating exploration, meaning proxy, and alignment, meaning primary, it restructures post-training into a modular, asynchronous, and reusable pipeline.
  • It supports weak-to-strong improvement by transferring relative improvement signals rather than the absolute distribution of the proxy.
  • Exploration signals from a small proxy such as Qwen3-1.7B improve the larger primary model in a controllable way.
  • Compared with GRPO, OPD-based distribution alignment converges within 70 steps, which demonstrates exploration efficiency.
  • The transfer strength can be controlled with a scaling factor, which gives robust performance across different settings.
  • Caching, reuse, and cross-model transfer of exploration signals greatly improve cost efficiency.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)