Freeform Preference Learning for Robotic Manipulation
- Published
- Source
- arXiv
- Paper number
- 553
- Field
- Robotics
- arXiv ID
- 2606.32027
Key points
- It resolves the ambiguity of binary preferences by defining natural-language axes such as speed, safety, and precision, and collecting pairwise preferences along each axis.
- It extends Bradley-Terry to multiple dimensions so that a single reward model conditioned on language labels outputs axis-specific scalar rewards.
- A VLM-based reward model using Qwen VL 3.5 4B combines visual and language understanding for reward learning.
- It achieves a 38-point gain over sparse reward and also a significant improvement over binary preference.
- It exhibits emergent compositionality, combining fast speed with new target combinations that were not present in the training data.
- By changing the reward condition at test time, it can steer policy behavior without retraining.
Paper links
External research summaries. These are not HDATF publications or measured product results.