TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
- Published
- Source
- arXiv
- Paper number
- 669
- Field
- Computer Vision
- arXiv ID
- 2607.17423
Key points
- It defines video evidence finding as a problem over a variable-size set of temporal segments rather than a single interval, so it can handle long videos, repeated scenes, first-person video, and question-form requests in one model.
- It solves data quality by stepwise verification instead of labeling everything at once, using candidate proposal, independent re-search, agent agreement, semantic verification, and boundary refinement.
- Its key technique is the temporal Wasserstein reward, which treats the predicted and ground-truth intervals as uniform distributions and computes the exact 1D W1 distance to give dense feedback without explicit matching.
- Overlap-based rewards such as tIoU give zero signal when the predicted interval does not overlap the ground truth at all, while Wasserstein reward distinguishes near misses from far misses in that case.
- Empirically, the fraction of groups with all equal rewards and no learning signal drops from 13.8% to 3.6%, and 75.8% of the groups that were all zero regain a meaningful ranking.
- The practical takeaway is that a much smaller model beats much larger ones, which is strong evidence that making answers temporally grounded and human-checkable works.
Paper links
External research summaries. These are not HDATF publications or measured product results.