Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
- Published
- Source
- arXiv
- Paper number
- 605
- Field
- Machine Learning
- arXiv ID
- 2607.11505
Key points
- By separating exploration, meaning proxy, and alignment, meaning primary, it restructures post-training into a modular, asynchronous, and reusable pipeline.
- It supports weak-to-strong improvement by transferring relative improvement signals rather than the absolute distribution of the proxy.
- Exploration signals from a small proxy such as Qwen3-1.7B improve the larger primary model in a controllable way.
- Compared with GRPO, OPD-based distribution alignment converges within 70 steps, which demonstrates exploration efficiency.
- The transfer strength can be controlled with a scaling factor, which gives robust performance across different settings.
- Caching, reuse, and cross-model transfer of exploration signals greatly improve cost efficiency.
Paper links
External research summaries. These are not HDATF publications or measured product results.