Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models
- Published
- Source
- arXiv
- Paper number
- 307
- Field
- LLMs / NLP
- arXiv ID
- 2606.05161
Key points
- The paper examines this question using an audio-fixed counterfactual method that removes only conflicting text and measures the resulting change in model preferences.
- Across five ALMs and four conflict tasks, sign flips appeared in 64.1 percent of conflict samples. In other words, the same audio-only branch preferred the audio-supported answer, while the combined branch preferred the text-supported answer.
- Under a strict 5 percentage point fidelity-drop budget, GACL improves nAUC by 17.8 points over the best contrastive-learning baseline and transfers to vision-text mediation without retraining, with up to a 40.5 percentage point gain.
Paper links
External research summaries. These are not HDATF publications or measured product results.