Projector Is All You Train
- Published
- Source
- arXiv
- Paper number
- 979
- Field
- LLMs / NLP
- arXiv ID
- 2608.19726
Key points
- In a 3D MLLM, freezing both the encoder and language-model backbone and training only the projector, a 3-layer MLP, still matched existing baselines on 3D classification and captioning.
- Under the same budget of 16 GPU-hours on a single A100, projector-only training performed similarly to jointly training the backbone with LoRA while delivering approximately 2× the training-sample throughput.
- Jointly training the backbone severely degraded existing capabilities: the Llama backbone's GSM8K score fell from 86.96 to 0.61 and HumanEval from 64.02 to 0.00. Projector-only training leaves the backbone untouched, inherently avoiding this regression.
- It analyzed apparent improvements in some vision scores after joint backbone training as artifacts of multiple-choice scoring, which renormalizes option probabilities, while instruction-following failures actually increased.
- It proposed a modular multimodal setup as a future direction, independently attaching and training modality-specific projectors on a single language-model backbone.
Paper links
External research summaries. These are not HDATF publications or measured product results.