Agent Priors-guided Policy Learning

Published
Source
arXiv
Paper number
1133
Field
Robotics
arXiv ID
2609.35690

Key points

  • Used each skill's structural prior both as an inductive bias during training and as a descriptive document for the runtime agent.
  • Proposed several priors per skill and trained and verified a separate diffusion policy for each prior.
  • Let the runtime agent read priors, handoff conditions, and verification evidence to choose policies and invoke them in chunks of up to 300 steps.
  • Showed stronger out-of-distribution skill generalization on MetaWorld and ManiSkill and succeeded at previously unseen skill compositions.
  • Confirmed the interface is essential: removing the prior information in ablations sharply degraded performance.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)