A Sober Look at the Robustness of CLIPs to Spurious Features

Wang, Qizhou; Lin, Yong; Chen, Yongqiang; Schmidt, Ludwig; Han, Bo; Zhang, Tong

Computer Science > Computer Vision and Pattern Recognition

arXiv:2403.11497 (cs)

[Submitted on 18 Mar 2024 (v1), last revised 2 Nov 2024 (this version, v2)]

Title:A Sober Look at the Robustness of CLIPs to Spurious Features

Authors:Qizhou Wang, Yong Lin, Yongqiang Chen, Ludwig Schmidt, Bo Han, Tong Zhang

View PDF HTML (experimental)

Abstract:Large vision language models, such as CLIP, demonstrate impressive robustness to spurious features than single-modal models trained on ImageNet. However, existing test datasets are typically curated based on ImageNet-trained models, which aim to capture the spurious features inherited in ImageNet. Benchmarking CLIP models based on the ImageNet-oriented spurious features may not be sufficient to reflect the extent to which CLIP models are robust to spurious correlations within CLIP training data, e.g., LAION. To this end, we craft a new challenging dataset named CounterAnimal designed to reveal the reliance of CLIP models on realistic spurious features. Specifically, we split animal photos into groups according to the backgrounds, and then identify a pair of groups for each class where a CLIP model shows high-performance drops across the two groups. Our evaluations show that the spurious features captured by CounterAnimal are generically learned by CLIP models with different backbones and pre-train data, yet have limited influence for ImageNet models. We provide theoretical insights that the CLIP objective cannot offer additional robustness. Furthermore, we also re-evaluate strategies such as scaling up parameters and high-quality pre-trained data. We find that they still help mitigate the spurious features, providing a promising path for future developments.

Comments:	NeurIPS 2024; Qizhou Wang, Yong Lin, and Yongqiang Chen contributed equally; Project page: this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:2403.11497 [cs.CV]
	(or arXiv:2403.11497v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2403.11497

Submission history

From: Yongqiang Chen [view email]
[v1] Mon, 18 Mar 2024 06:04:02 UTC (1,777 KB)
[v2] Sat, 2 Nov 2024 04:58:15 UTC (2,396 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:A Sober Look at the Robustness of CLIPs to Spurious Features

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:A Sober Look at the Robustness of CLIPs to Spurious Features

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators