$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 1024
- Field
- Robotics
- arXiv ID
- 2608.26053
Key points
- Unlike prior approaches that used language reasoning only as an auxiliary training signal, it generates reasoning at test time to direct a lower-level policy.
- Stage 1 initializes the reasoning style from expert-reasoner traces (regardless of success or failure), and stage 2 performs offline RL (Dr. GRPO) using rubric-based rewards from a VLM judge.
- Across 12 held-out grocery-packing tasks, it achieved an average success rate of 47.9%, outperforming instruction-only imitation learning (38.0%).
- The trained reasoner exhibited deliberative, action-oriented reasoning, such as comparing alternatives, observing the scene again when uncertain, and choosing steps incrementally.
- Intervention experiments that adjusted the reasoning budget confirmed that reasoning contributes causally to performance.
Paper links
External research summaries. These are not HDATF publications or measured product results.