ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning
- Published
- Source
- arXiv
- Paper number
- 288
- Field
- Computer Vision
- arXiv ID
- 2606.02576
Key points
- Multimodal large language models achieve strong performance through instruction tuning, but real-world deployment requires them to keep acquiring new vision-language abilities, so multimodal continual instruction tuning is essential.
- To address this problem, the paper proposes ProtoAda, a prototype-guided adaptive tuning framework.
- Extensive experiments across multiple benchmarks show that ProtoAda achieves superior performance, especially on tasks where answer structure is easily damaged by sequential tuning.
Paper links
External research summaries. These are not HDATF publications or measured product results.