ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning

Published
Source
arXiv
Paper number
288
Field
Computer Vision
arXiv ID
2606.02576

Key points

  • Multimodal large language models achieve strong performance through instruction tuning, but real-world deployment requires them to keep acquiring new vision-language abilities, so multimodal continual instruction tuning is essential.
  • To address this problem, the paper proposes ProtoAda, a prototype-guided adaptive tuning framework.
  • Extensive experiments across multiple benchmarks show that ProtoAda achieves superior performance, especially on tasks where answer structure is easily damaged by sequential tuning.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)