Cambrian-P: Pose-Grounded Video Understanding
- Published
- Source
- arXiv
- Paper number
- 197
- Field
- Computer Vision
- arXiv ID
- 2605.22819
Key points
- Annotation: use the off-the-shelf pose estimator ViPE to generate external and internal camera parameters for the remaining clips.
- Model size: as the LLM size increases from 0.5B to 7B parameters, both VQA performance and pose accuracy improve consistently.
- Data size: increasing the amount of training data widens the performance gap between the pose-grounding model and the baseline.
Paper links
External research summaries. These are not HDATF publications or measured product results.