Agent Priors-guided Policy Learning
- Published
- Source
- arXiv
- Paper number
- 1133
- Field
- Robotics
- arXiv ID
- 2609.35690
Key points
- Used each skill's structural prior both as an inductive bias during training and as a descriptive document for the runtime agent.
- Proposed several priors per skill and trained and verified a separate diffusion policy for each prior.
- Let the runtime agent read priors, handoff conditions, and verification evidence to choose policies and invoke them in chunks of up to 300 steps.
- Showed stronger out-of-distribution skill generalization on MetaWorld and ManiSkill and succeeded at previously unseen skill compositions.
- Confirmed the interface is essential: removing the prior information in ablations sharply degraded performance.
Paper links
External research summaries. These are not HDATF publications or measured product results.