SupMAE: Supervised Masked Autoencoders Are Efficient Vision Learners

Liang, Feng; Li, Yangguang; Marculescu, Diana

Computer Science > Computer Vision and Pattern Recognition

arXiv:2205.14540 (cs)

[Submitted on 28 May 2022 (v1), last revised 21 Jan 2024 (this version, v3)]

Title:SupMAE: Supervised Masked Autoencoders Are Efficient Vision Learners

Authors:Feng Liang, Yangguang Li, Diana Marculescu

View PDF HTML (experimental)

Abstract:Recently, self-supervised Masked Autoencoders (MAE) have attracted unprecedented attention for their impressive representation learning ability. However, the pretext task, Masked Image Modeling (MIM), reconstructs the missing local patches, lacking the global understanding of the image. This paper extends MAE to a fully supervised setting by adding a supervised classification branch, thereby enabling MAE to learn global features from golden labels effectively. The proposed Supervised MAE (SupMAE) only exploits a visible subset of image patches for classification, unlike the standard supervised pre-training where all image patches are used. Through experiments, we demonstrate that SupMAE is not only more training efficient but it also learns more robust and transferable features. Specifically, SupMAE achieves comparable performance with MAE using only 30% of compute when evaluated on ImageNet with the ViT-B/16 model. SupMAE's robustness on ImageNet variants and transfer learning performance outperforms MAE and standard supervised pre-training counterparts. Codes are available at this https URL.

Comments:	Edge Intelligence Workshop Workshop at AAAI 2024
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as:	arXiv:2205.14540 [cs.CV]
	(or arXiv:2205.14540v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2205.14540

Submission history

From: Feng Liang [view email]
[v1] Sat, 28 May 2022 23:05:03 UTC (1,488 KB)
[v2] Tue, 16 Aug 2022 17:49:32 UTC (512 KB)
[v3] Sun, 21 Jan 2024 02:12:04 UTC (859 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:SupMAE: Supervised Masked Autoencoders Are Efficient Vision Learners

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:SupMAE: Supervised Masked Autoencoders Are Efficient Vision Learners

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators