Déjà View: Looping Transformers for Multi-View 3D Reconstruction
- Published
- Source
- arXiv
- Paper number
- 271
- Field
- Computer Vision
- arXiv ID
- 2605.30215
Key points
- Register tokens are four tokens used as a global scratchpad for the transformer.
- Camera tokens are special tokens that store information about the pose and intrinsics of the current view.
- Frame attention is a subblock that operates independently on each image, allowing the model to infer local geometry within a single view.
Paper links
External research summaries. These are not HDATF publications or measured product results.