Freeform Preference Learning for Robotic Manipulation

Published
Source
arXiv
Paper number
553
Field
Robotics
arXiv ID
2606.32027

Key points

  • It resolves the ambiguity of binary preferences by defining natural-language axes such as speed, safety, and precision, and collecting pairwise preferences along each axis.
  • It extends Bradley-Terry to multiple dimensions so that a single reward model conditioned on language labels outputs axis-specific scalar rewards.
  • A VLM-based reward model using Qwen VL 3.5 4B combines visual and language understanding for reward learning.
  • It achieves a 38-point gain over sparse reward and also a significant improvement over binary preference.
  • It exhibits emergent compositionality, combining fast speed with new target combinations that were not present in the training data.
  • By changing the reward condition at test time, it can steer policy behavior without retraining.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)