RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
- Published
- Source
- arXiv
- Paper number
- 665
- Field
- Robotics
- arXiv ID
- 2607.17977
Key points
- The core claim is that it is not enough to understand vision well; the representation and output must be immediately usable for robot manipulation.
- As a method, it adds contact-point prediction, meaning a grasp center and gripper-plane rotation angle, and native 3D localization as pretraining tasks, and 3D localization is included in the 2B and 9B models.
- The architecture is a decoder-only vision-language model based on Qwen3.5, and it uses DeepStack and Interleaved MRoPE to handle images, multi-view images, and video over long contexts.
- To unify multiple robots under one policy, it uses a shared action space aligned by body part and masks only the axes that each robot supports.
- As a result, the average success rate over three real-robot tasks is 86.67 percent for RynnBrain-VLA, 73.33 percent for GR00T N1.7, 65.00 percent for π0.5, and 60.00 percent for the Qwen-based model.
- The generalist version trained on multiple tasks and multiple robots performs better than task-specific training, with success rising from 86.67 percent to 91.67 percent, which supports the idea that heterogeneous robot data can complement one another.
Paper links
External research summaries. These are not HDATF publications or measured product results.