DenseReward: Dense Reward Learning via Failure Synthesis for Robotic Manipulation

Published
Source
arXiv
Paper number
627
Field
Robotics
arXiv ID
2607.13033

Key points

  • It automatically synthesized trajectories of physical failures such as collisions, failed grasps, dropped objects, and recovery in simulation without human labels.
  • It trained a dense reward model that takes visual observations and language instructions and scores task progress at every frame.
  • In both simulated and real manipulation, it predicted dense rewards more accurately than general-purpose VLMs and existing robot reward models.
  • These rewards provide fine-grained progress signals for model-predictive control and reinforcement learning and can be used to improve robot policies after imitation learning.
  • The current scope focuses on relatively short manipulation tasks, leaving tool use, long-horizon work, and rewards reflecting human preferences for future work.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)