SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
- Published
- Source
- arXiv
- Paper number
- 259
- Field
- Computer Vision
- arXiv ID
- 2605.27367
Key points
- DENSE focuses on long-time-series video sequences and feeds in nearly every frame up to the 500-frame limit.
- For robotics, use models trained on DA-Next-5M or similar first-person datasets; standard indoor and outdoor foundation models are likely to fail on wrist-mounted manipulation tasks.
- For AR and VR, streaming or chunk-based models are preferred because they must handle long sessions without exceeding the memory limits of mobile devices or headsets.
Paper links
External research summaries. These are not HDATF publications or measured product results.