Vector Policy Optimization: Training for Diversity Improves Test-Time Search

Published
Source
arXiv
Paper number
206
Field
Machine Learning
arXiv ID
2605.22817

Key points

  • Maze Navigation is a synthetic task in which an agent must collect items such as gold and diamonds while avoiding lava hazards. Because the goals are deliberately conflicting, for example, gold is often near lava, it is a perfect testbed for validating Pareto-front coverage.
  • MuSiQue is a multi-hop QA task that rewards both final answer accuracy and citation accuracy for intermediate reasoning steps.
  • ToolRL is a function-calling benchmark in which the model must satisfy structural constraints while also achieving high attribute-level accuracy.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)