Robustifying Vision-Language Models via Test-Time Prompt Adaptation
- Published
- Source
- arXiv
- Paper number
- 615
- Field
- Computer Vision
- arXiv ID
- 2607.09450
Key points
- RITA reduces adversarial outliers by using optimal transport to align the feature distribution of multiple augmentations of a test image with the text-prototype distribution.
- A dynamic cache accumulates reliable cues from the test stream to continually refine prompts online.
- Across multiple experiments, it improved VLM robustness to adversarial attacks without reducing clean-image accuracy.
- Because it requires neither model retraining nor access to labeled data, it can be added to existing CLIP-family classification systems as a test-time defense.
- Current validation centers on image classification. Extension to generative tasks such as image captioning remains open, and the method does not eliminate every risk from new attacks.
Paper links
External research summaries. These are not HDATF publications or measured product results.