SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation

Luo, Huaishao; Bao, Junwei; Wu, Youzheng; He, Xiaodong; Li, Tianrui

Computer Science > Computer Vision and Pattern Recognition

arXiv:2211.14813v2 (cs)

[Submitted on 27 Nov 2022 (v1), last revised 20 Jun 2023 (this version, v2)]

Title:SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation

Authors:Huaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He, Tianrui Li

View PDF

Abstract:Recently, the contrastive language-image pre-training, e.g., CLIP, has demonstrated promising results on various downstream tasks. The pre-trained model can capture enriched visual concepts for images by learning from a large scale of text-image data. However, transferring the learned visual knowledge to open-vocabulary semantic segmentation is still under-explored. In this paper, we propose a CLIP-based model named SegCLIP for the topic of open-vocabulary segmentation in an annotation-free manner. The SegCLIP achieves segmentation based on ViT and the main idea is to gather patches with learnable centers to semantic regions through training on text-image pairs. The gathering operation can dynamically capture the semantic groups, which can be used to generate the final segmentation results. We further propose a reconstruction loss on masked patches and a superpixel-based KL loss with pseudo-labels to enhance the visual representation. Experimental results show that our model achieves comparable or superior segmentation accuracy on the PASCAL VOC 2012 (+0.3% mIoU), PASCAL Context (+2.3% mIoU), and COCO (+2.2% mIoU) compared with baselines. We release the code at this https URL.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2211.14813 [cs.CV]
	(or arXiv:2211.14813v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2211.14813

Submission history

From: Huaishao Luo [view email]
[v1] Sun, 27 Nov 2022 12:38:52 UTC (1,110 KB)
[v2] Tue, 20 Jun 2023 06:36:09 UTC (1,124 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators