GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
- Published
- Source
- arXiv
- Paper number
- 935
- Field
- Robotics
- arXiv ID
- 2608.15875
Key points
- System 2 uses PaliGemma2 3B to understand the scene and task, System 1 uses PaliGemma2 3B plus a 0.5B action model to produce control values, and System 3 uses a 5B world model to predict future video and progress.
- The trajectory data consists of 20,535.65 hours of real-robot data, 8,251.83 hours of UMI human demonstrations, 2,862.36 hours of first-person human video, 1,453.92 hours of simulation, and 4,153.22 hours of world-model generation.
- In the System 3 ablation, the gift-wrapping success rate on AgileX PiPER-X rises from 0 percent for the base model to 80 percent when future video and progress are used together, and the clothing-folding completion time on AgileX PiPER falls from 107 seconds to 75 seconds.
- Under the Co-Train setting on 50 RoboTwin 2.0 tasks, where 50 clean demonstrations per task and one policy are jointly trained, the results are 66.8 percent for Easy, 67.9 percent for Hard, and 67.35 percent on average, while pi0.5 reaches 70.7 percent, 46.0 percent, and 58.35 percent.
- The real-robot tasks are usually evaluated with only 10 to 20 runs per setting and no confidence intervals, and the final MiMo score of 0.5704 also uses evaluation data that was included in training, so it is not an independent test result.
Paper links
External research summaries. These are not HDATF publications or measured product results.