The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs

Published
Source
arXiv
Paper number
602
Field
Computer Vision
arXiv ID
2607.09544

Key points

  • It trains three probes over VLM internal activations to predict ground-truth count, output count, and whether an error is present, and the probes reliably predict mistakes.
  • SVCCA shows that the correct-count probe and the output-count probe share a partial subspace, but the readout is misaligned, which is a structural cause of error.
  • Causal steering confirms the cause, because strengthening the count direction improves counting performance while random directions hurt it.
  • Detector-guided self-correction applies selective reprompting only when an error is predicted, which improves accuracy by up to 15.6 absolute points.
  • It improves performance through inference-time intervention alone, so the method requires no parameter update and is practical to deploy.
  • There is a strong positive correlation, with rho equal to 0.803, between probe F1 and correction gains, which means that internal error-signal quality determines correction effectiveness.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)