Déjà View: Looping Transformers for Multi-View 3D Reconstruction

Published
Source
arXiv
Paper number
271
Field
Computer Vision
arXiv ID
2605.30215

Key points

  • Register tokens are four tokens used as a global scratchpad for the transformer.
  • Camera tokens are special tokens that store information about the pose and intrinsics of the current view.
  • Frame attention is a subblock that operates independently on each image, allowing the model to infer local geometry within a single view.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)