Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
- Published
- Source
- arXiv
- Paper number
- 451
- Field
- LLMs / NLP
- arXiv ID
- 2606.18216
Key points
- It addresses both the mode-seeking bias of teacher-logit imitation and the drift problem caused by injecting teacher response gradients.
- BCQ places teacher-correct and student-wrong answers as anonymous candidates so the student has to discriminate them while staying on-policy.
- NCQ aggregates the student's failed rollouts and presents common failure patterns in the prompt.
- A prompt replay buffer recycles hard questions to amplify learning within the student's zone of proximal development.
- At 0.8B, it improves average performance by +4.9 points for VLMs, +4.4 points for LLMs, and +2.3 points for video models compared with GRPO, with a large +16.9-point gain on GPQA-D.
- It confirms that the gains from ZPPO grow with teacher size and shrink with 4B and 9B teachers.
Paper links
External research summaries. These are not HDATF publications or measured product results.