Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling

Published
Source
arXiv
Paper number
286
Field
Computer Vision
arXiv ID
2606.02578

Key points

  • To address this issue, the paper introduces the Perceptually Perturbed Judgment Dataset, which constructs minimally edited counterfactual responses that separate perceptual errors and enable verifiable supervision.
  • Experiments across multiple MLLM-as-a-Judge benchmarks show that this approach substantially improves perceptual fidelity, ranking consistency, and alignment with human judgments.
  • The result establishes a scalable and generalizable path toward training multimodal judges that are perceptually grounded, interpretable, and robust to vision reasoning conflicts.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)