$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning

Published
Source
arXiv
Paper number
1024
Field
Robotics
arXiv ID
2608.26053

Key points

  • Unlike prior approaches that used language reasoning only as an auxiliary training signal, it generates reasoning at test time to direct a lower-level policy.
  • Stage 1 initializes the reasoning style from expert-reasoner traces (regardless of success or failure), and stage 2 performs offline RL (Dr. GRPO) using rubric-based rewards from a VLM judge.
  • Across 12 held-out grocery-packing tasks, it achieved an average success rate of 47.9%, outperforming instruction-only imitation learning (38.0%).
  • The trained reasoner exhibited deliberative, action-oriented reasoning, such as comparing alternatives, observing the scene again when uncertain, and choosing steps incrementally.
  • Intervention experiments that adjusted the reasoning budget confirmed that reasoning contributes causally to performance.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)