Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation

Richardson, Elad; Alaluf, Yuval; Patashnik, Or; Nitzan, Yotam; Azar, Yaniv; Shapiro, Stav; Cohen-Or, Daniel

Computer Science > Computer Vision and Pattern Recognition

arXiv:2008.00951v2 (cs)

[Submitted on 3 Aug 2020 (v1), last revised 21 Apr 2021 (this version, v2)]

Title:Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation

Authors:Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, Daniel Cohen-Or

View PDF

Abstract:We present a generic image-to-image translation framework, pixel2style2pixel (pSp). Our pSp framework is based on a novel encoder network that directly generates a series of style vectors which are fed into a pretrained StyleGAN generator, forming the extended W+ latent space. We first show that our encoder can directly embed real images into W+, with no additional optimization. Next, we propose utilizing our encoder to directly solve image-to-image translation tasks, defining them as encoding problems from some input domain into the latent domain. By deviating from the standard invert first, edit later methodology used with previous StyleGAN encoders, our approach can handle a variety of tasks even when the input image is not represented in the StyleGAN domain. We show that solving translation tasks through StyleGAN significantly simplifies the training process, as no adversary is required, has better support for solving tasks without pixel-to-pixel correspondence, and inherently supports multi-modal synthesis via the resampling of styles. Finally, we demonstrate the potential of our framework on a variety of facial image-to-image translation tasks, even when compared to state-of-the-art solutions designed specifically for a single task, and further show that it can be extended beyond the human facial domain.

Comments:	Accepted to CVPR 2021, project page available at this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2008.00951 [cs.CV]
	(or arXiv:2008.00951v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2008.00951

Submission history

From: Elad Richardson [view email]
[v1] Mon, 3 Aug 2020 15:30:38 UTC (13,561 KB)
[v2] Wed, 21 Apr 2021 12:53:36 UTC (33,056 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators