Cambrian-P: Pose-Grounded Video Understanding

Published
Source
arXiv
Paper number
197
Field
Computer Vision
arXiv ID
2605.22819

Key points

  • Annotation: use the off-the-shelf pose estimator ViPE to generate external and internal camera parameters for the remaining clips.
  • Model size: as the LLM size increases from 0.5B to 7B parameters, both VQA performance and pose accuracy improve consistently.
  • Data size: increasing the amount of training data widens the performance gap between the pose-grounding model and the baseline.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)