G0.5: One Autoregressive Stream for Robot Reasoning and Action
- Published
- Source
- arXiv
- Paper number
- 887
- Field
- Robotics
- arXiv ID
- 2608.11739
Key points
- Reasoning and action use the same model weights and the same token stream, and there is no separate action-expert model.
- It trains 14 robot embodiments in both real and simulated environments and maps their different action spaces into one 27-dimensional shared space and one token vocabulary.
- The thought process includes subgoals, object locations, 2D movement paths, and action hints, while visual memory uses the most recent few seconds of video.
- The average success rates are 76.7 percent on real robots, 82.5 percent on DROID's new environments and objects, 98.9 percent on LIBERO, 93.3 percent on RoboTwin 2.0, and 87.3 percent on SimplerEnv-Bridge.
- The BEHAVIOR long-horizon score is 31.4 percent, which is higher than the comparison model's 26.3 percent, but visual memory is limited to a few seconds and lower-body actions are not evaluated separately.
Paper links
External research summaries. These are not HDATF publications or measured product results.