Research
Robustness of Vision Foundation Models to Common Perturbations
arXiv:2604.14973v1 Announce Type: cross Abstract: A vision foundation model outputs an embedding vector for an image, which can be affected by common editing operations (e.g., JPEG compression, bright
arXiv:2604.14973v1 Announce Type: cross Abstract: A vision foundation model outputs an embedding vector for an image, which can be affected by common editing operations (e.g., JPEG compression, brightness, contrast adjustments). These common perturbations alter embedding vectors and may impact the performance of downstream tasks using these embeddings. In this work, we present the first systematic study on foundation models' robustness to such perturbations. We propose three robustness metrics and formulate five desired mathematical properties for these metrics, analyzing which properties they satisfy or violate. Using these metrics, we evaluate six industry-scale foundation models (OpenAI, Meta) across nine common perturbation categories, finding them generally non-robust. We also show that common perturbations degrade downstream application performance (e.g., classification accuracy) and that robustness values can predict performance impacts. Finally, we propose a fine-tuning approach to improve robustness without sacrificing utility.
Related
- EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model
- HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
- FF3R: Feedforward Feature 3D Reconstruction from Unconstrained views
- Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
- Do Vision Language Models Need to Process Image Tokens?
Source: arXiv cs.CV | 2026-04-17