IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

Published
Source
arXiv
Paper number
689
Field
Computer Vision
arXiv ID
2607.19228

Key points

  • It is the first streaming transformer that predicts geometry and instance identity at the same time from a video stream.
  • It processes frames sequentially with causal masking and reuses past context through a KV cache.
  • The authors build InsScene4D-147K, a 4D instance-annotated dataset with 147,000 scenes.
  • It shows much stronger performance on object tracking, with T-mIoU of 58.84 compared with 45.38 for SAM2.
  • Streaming clustering keeps memory constant at 0.7 GB regardless of the number of frames.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)