Projector Is All You Train

Published
Source
arXiv
Paper number
979
Field
LLMs / NLP
arXiv ID
2608.19726

Key points

  • In a 3D MLLM, freezing both the encoder and language-model backbone and training only the projector, a 3-layer MLP, still matched existing baselines on 3D classification and captioning.
  • Under the same budget of 16 GPU-hours on a single A100, projector-only training performed similarly to jointly training the backbone with LoRA while delivering approximately 2× the training-sample throughput.
  • Jointly training the backbone severely degraded existing capabilities: the Llama backbone's GSM8K score fell from 86.96 to 0.61 and HumanEval from 64.02 to 0.00. Projector-only training leaves the backbone untouched, inherently avoiding this regression.
  • It analyzed apparent improvements in some vision scores after joint backbone training as artifacts of multiple-choice scoring, which renormalizes option probabilities, while instruction-following failures actually increased.
  • It proposed a modular multimodal setup as a future direction, independently attaching and training modality-specific projectors on a single language-model backbone.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)