SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

Published
Source
arXiv
Paper number
259
Field
Computer Vision
arXiv ID
2605.27367

Key points

  • DENSE focuses on long-time-series video sequences and feeds in nearly every frame up to the 500-frame limit.
  • For robotics, use models trained on DA-Next-5M or similar first-person datasets; standard indoor and outdoor foundation models are likely to fail on wrist-mounted manipulation tasks.
  • For AR and VR, streaming or chunk-based models are preferred because they must handle long sessions without exceeding the memory limits of mobile devices or headsets.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)