GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

Published
Source
arXiv
Paper number
935
Field
Robotics
arXiv ID
2608.15875

Key points

  • System 2 uses PaliGemma2 3B to understand the scene and task, System 1 uses PaliGemma2 3B plus a 0.5B action model to produce control values, and System 3 uses a 5B world model to predict future video and progress.
  • The trajectory data consists of 20,535.65 hours of real-robot data, 8,251.83 hours of UMI human demonstrations, 2,862.36 hours of first-person human video, 1,453.92 hours of simulation, and 4,153.22 hours of world-model generation.
  • In the System 3 ablation, the gift-wrapping success rate on AgileX PiPER-X rises from 0 percent for the base model to 80 percent when future video and progress are used together, and the clothing-folding completion time on AgileX PiPER falls from 107 seconds to 75 seconds.
  • Under the Co-Train setting on 50 RoboTwin 2.0 tasks, where 50 clean demonstrations per task and one policy are jointly trained, the results are 66.8 percent for Easy, 67.9 percent for Hard, and 67.35 percent on average, while pi0.5 reaches 70.7 percent, 46.0 percent, and 58.35 percent.
  • The real-robot tasks are usually evaluated with only 10 to 20 runs per setting and no confidence intervals, and the final MiMo score of 0.5704 also uses evaluation data that was included in training, so it is not an independent test result.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)