Exploiting Diffusion Prior for Generalizable Dense Prediction

Lee, Hsin-Ying; Tseng, Hung-Yu; Lee, Hsin-Ying; Yang, Ming-Hsuan

Computer Science > Computer Vision and Pattern Recognition

arXiv:2311.18832 (cs)

[Submitted on 30 Nov 2023 (v1), last revised 2 Apr 2024 (this version, v2)]

Title:Exploiting Diffusion Prior for Generalizable Dense Prediction

Authors:Hsin-Ying Lee, Hung-Yu Tseng, Hsin-Ying Lee, Ming-Hsuan Yang

View PDF HTML (experimental)

Abstract:Contents generated by recent advanced Text-to-Image (T2I) diffusion models are sometimes too imaginative for existing off-the-shelf dense predictors to estimate due to the immitigable domain gap. We introduce DMP, a pipeline utilizing pre-trained T2I models as a prior for dense prediction tasks. To address the misalignment between deterministic prediction tasks and stochastic T2I models, we reformulate the diffusion process through a sequence of interpolations, establishing a deterministic mapping between input RGB images and output prediction distributions. To preserve generalizability, we use low-rank adaptation to fine-tune pre-trained models. Extensive experiments across five tasks, including 3D property estimation, semantic segmentation, and intrinsic image decomposition, showcase the efficacy of the proposed method. Despite limited-domain training data, the approach yields faithful estimations for arbitrary images, surpassing existing state-of-the-art algorithms.

Comments:	To appear in CVPR 2024. Project page: this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2311.18832 [cs.CV]
	(or arXiv:2311.18832v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2311.18832

Submission history

From: Hsin-Ying Lee [view email]
[v1] Thu, 30 Nov 2023 18:59:44 UTC (7,746 KB)
[v2] Tue, 2 Apr 2024 17:59:33 UTC (17,451 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Exploiting Diffusion Prior for Generalizable Dense Prediction

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Exploiting Diffusion Prior for Generalizable Dense Prediction

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators