The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs
- Published
- Source
- arXiv
- Paper number
- 602
- Field
- Computer Vision
- arXiv ID
- 2607.09544
Key points
- It trains three probes over VLM internal activations to predict ground-truth count, output count, and whether an error is present, and the probes reliably predict mistakes.
- SVCCA shows that the correct-count probe and the output-count probe share a partial subspace, but the readout is misaligned, which is a structural cause of error.
- Causal steering confirms the cause, because strengthening the count direction improves counting performance while random directions hurt it.
- Detector-guided self-correction applies selective reprompting only when an error is predicted, which improves accuracy by up to 15.6 absolute points.
- It improves performance through inference-time intervention alone, so the method requires no parameter update and is practical to deploy.
- There is a strong positive correlation, with rho equal to 0.803, between probe F1 and correction gains, which means that internal error-signal quality determines correction effectiveness.
Paper links
External research summaries. These are not HDATF publications or measured product results.