Vector Policy Optimization: Training for Diversity Improves Test-Time Search
- Published
- Source
- arXiv
- Paper number
- 206
- Field
- Machine Learning
- arXiv ID
- 2605.22817
Key points
- Maze Navigation is a synthetic task in which an agent must collect items such as gold and diamonds while avoiding lava hazards. Because the goals are deliberately conflicting, for example, gold is often near lava, it is a perfect testbed for validating Pareto-front coverage.
- MuSiQue is a multi-hop QA task that rewards both final answer accuracy and citation accuracy for intermediate reasoning steps.
- ToolRL is a function-calling benchmark in which the model must satisfy structural constraints while also achieving high attribute-level accuracy.
Paper links
External research summaries. These are not HDATF publications or measured product results.