ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

Published
Source
arXiv
Paper number
816
Field
Computer Vision
arXiv ID
2608.04010

Key points

  • It proposes the first framework that creates separate parallel branches for ViT and LLM so that compute allocation can be adjusted independently.
  • A prefix-conditioned branch shares backbone parameters, so compute can be expanded without increasing parameters.
  • Systematic experiments on nine vision-language branch configurations show that the optimal allocation differs by task.
  • It confirms consistent gains over a single-branch model at 1B, 2B, and 8B scale.
  • It analyzes that parallel execution does not directly translate extra compute into wall-clock latency.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)