Lance: Unified Multimodal Modeling by Multi-Task Synergy

Published
Source
arXiv
Paper number
196
Field
Computer Vision
arXiv ID
2605.18678

Key points

  • For understanding, generation helps: training on generative tasks such as video editing makes the model develop a deeper understanding of temporal dynamics and spatial relationships, which improves performance on video understanding benchmarks.
  • For video generation, Lance scores 85.11 on VBench, the highest among unified models at the time of publication, with strong performance on object grounding and temporal consistency.
  • For multimodal understanding, Lance scores 62.0 on MVBench, which evaluates temporal cognition and video reasoning. This is an 11.3% improvement over the next-best unified model, Show-o2 (7B).

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)