IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer
- Published
- Source
- arXiv
- Paper number
- 689
- Field
- Computer Vision
- arXiv ID
- 2607.19228
Key points
- It is the first streaming transformer that predicts geometry and instance identity at the same time from a video stream.
- It processes frames sequentially with causal masking and reuses past context through a KV cache.
- The authors build InsScene4D-147K, a 4D instance-annotated dataset with 147,000 scenes.
- It shows much stronger performance on object tracking, with T-mIoU of 58.84 compared with 45.38 for SAM2.
- Streaming clustering keeps memory constant at 0.7 GB regardless of the number of frames.
Paper links
External research summaries. These are not HDATF publications or measured product results.