Lance: Unified Multimodal Modeling by Multi-Task Synergy
- Published
- Source
- arXiv
- Paper number
- 196
- Field
- Computer Vision
- arXiv ID
- 2605.18678
Key points
- For understanding, generation helps: training on generative tasks such as video editing makes the model develop a deeper understanding of temporal dynamics and spatial relationships, which improves performance on video understanding benchmarks.
- For video generation, Lance scores 85.11 on VBench, the highest among unified models at the time of publication, with strong performance on object grounding and temporal consistency.
- For multimodal understanding, Lance scores 62.0 on MVBench, which evaluates temporal cognition and video reasoning. This is an 11.3% improvement over the next-best unified model, Show-o2 (7B).
Paper links
External research summaries. These are not HDATF publications or measured product results.