DenseReward: Dense Reward Learning via Failure Synthesis for Robotic Manipulation
- Published
- Source
- arXiv
- Paper number
- 627
- Field
- Robotics
- arXiv ID
- 2607.13033
Key points
- It automatically synthesized trajectories of physical failures such as collisions, failed grasps, dropped objects, and recovery in simulation without human labels.
- It trained a dense reward model that takes visual observations and language instructions and scores task progress at every frame.
- In both simulated and real manipulation, it predicted dense rewards more accurately than general-purpose VLMs and existing robot reward models.
- These rewards provide fine-grained progress signals for model-predictive control and reinforcement learning and can be used to improve robot policies after imitation learning.
- The current scope focuses on relatively short manipulation tasks, leaving tool use, long-horizon work, and rewards reflecting human preferences for future work.
Paper links
External research summaries. These are not HDATF publications or measured product results.