Fetching the paper…
Reading the bibliography…
Can a generative model be trained to produce images from a specific domain, guided by a text prompt only, without seeing any image? In other words: can an image generator be trained "blindly"? Leveraging the semantic power of large scale Contrastive-Language-Image-Pre-training (CLIP) models, we present a text-driven method that allows shifting a generative model to new domains, without having to collect even a single image.
On lines and planes of closest fit to systems of points in space
Karl Pearson F.R.S · 1901
Earlier work this paper cites.
John Wiley and Sons, Ltd, 1990
Partitioning Around Medoids (Program PAM) · 1990
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop
Fisher Yu, Ari Seff, Yinda Zhang, Shuran Song, Thomas Funkhouser, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Photo-realistic single image super-resolution using a generative adversarial network
Christian Ledig, Lucas Theis, Ferenc Huszár, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Large scale gan training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2018
Earlier work this paper cites.
Coco-stuff: Thing and stuff classes in context
Holger Caesar, Jasper Uijlings, and Vittorio Ferrari · 2018
Earlier work this paper cites.
Deep network interpolation for continuous imagery effect transition, 2018
Xintao Wang, Ke Yu, Chao Dong, Xiaoou Tang, and Chen Change Loy · 2018
Earlier work this paper cites.
Transferring gans: generating images from limited data
Yaxing Wang, Chenshen Wu, Luis Herranz, Joost van de Weijer, Abel Gonzalez-Garcia, and B. Raducanu · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
Image2stylegan: How to embed images into the stylegan latent space?
Rameen Abdal, Yipeng Qin, and Peter Wonka · 2019
Earlier work this paper cites.
scikit-learn-extra module, 2019
Christos Aridas, Jan-Oliver Joswig, Timothée Mathieu, and Roman Yurchak · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, C. Hsieh, and Kai-Wei Chang · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, D. Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Image generation from small datasets via batch statistics adaptation
Atsuhiro Noguchi and Tatsuya Harada · 2019
Earlier work this paper cites.
Semantic image synthesis with spatially-adaptive normalization
Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu · 2019
Earlier work this paper cites.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Hao Tan and Mohit Bansal · 2019
Earlier work this paper cites.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, A. E. Kholy, Faisal Ahmed, Zhe Gan, Y. Cheng, and Jing jing Liu · 2020
Earlier work this paper cites.
Stargan v2: Diverse image synthesis for multiple domains
Yunjey Choi, Youngjung Uh, Jaejun Yoo, and Jung-Woo Ha · 2020
Earlier work this paper cites.
VirTex: Learning visual representations from textual annotations
Karan Desai and J. Johnson · 2020
Cited alongside, same era.
Ganspace: Discovering interpretable gan controls
Erik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, and Sylvain Paris · 2020
Cited alongside, same era.
Training generative adversarial networks with limited data
Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila · 2020
Cited alongside, same era.
Analyzing and improving the image quality of stylegan
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila · 2020
Cited alongside, same era.
Unicoder-VL: A universal encoder for vision and language by cross-modal pre-training
Gen Li, N. Duan, Yuejian Fang, Daxin Jiang, and M. Zhou · 2020
Cited alongside, same era.
Paint by word, 2021
David Bau, Alex Andonian, Audrey Cui, YeonHwan Park, Ali Jahanian, Aude Oliva, and Antonio Torralba · 2021
Closest in time.
Vqgan + clip, 2021
Katherine Crowson · 2021
Closest in time.
Deep neural networks are surprisingly reversible: A baseline for zero-shot inversion, 2021
Xin Dong, Hongxu Yin, Jose M. Alvarez, Jan Kautz, and Pavlo Molchanov · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Alias-free generative adversarial networks, 2021
Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila · 2021
Closest in time.
The big sleep, 2021
Ryan Murdock · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, C. Li, X. Hu, Pengchuan Zhang, Lei Zhang, Longguang Wang, H. Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao · 2020
Cited alongside, same era.
Few-shot image generation with elastic weight consolidation
Yijun Li, Richard Zhang, Jingwan Lu, and Eli Shechtman · 2020
Cited alongside, same era.
Towards faster and stabilized gan training for high-fidelity few-shot image synthesis
Bingchen Liu, Yizhe Zhu, Kunpeng Song, and Ahmed Elgammal · 2020
Cited alongside, same era.
Freeze the discriminator: a simple baseline for fine-tuning gans
Sangwoo Mo, Minsu Cho, and Jinwoo Shin · 2020
Cited alongside, same era.
Resolution dependent gan interpolation for controllable image synthesis between domains
Justin NM Pinkney and Doron Adler · 2020
Cited alongside, same era.
Encoding in style: a stylegan encoder for image-to-image translation
Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or · 2020
Cited alongside, same era.
Few-shot adaptation of generative adversarial networks
Esther Robb, Wen-Sheng Chu, Abhishek Kumar, and Jia-Bin Huang · 2020
Cited alongside, same era.
Large: Latent-based regression through gan semantics, 2021
Yotam Nitzan, Rinon Gal, Ofir Brenner, and Daniel Cohen-Or · 2021
Closest in time.
Few-shot image generation via cross-domain correspondence
Utkarsh Ojha, Yijun Li, Jingwan Lu, Alexei A Efros, Yong Jae Lee, Eli Shechtman, and Richard Zhang · 2021
Closest in time.
Styleclip: Text-driven manipulation of stylegan imagery
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Closest in time.
You only need adversarial supervision for semantic image synthesis
Edgar Schönfeld, Vadim Sushko, Dan Zhang, Juergen Gall, Bernt Schiele, and Anna Khoreva · 2021
Closest in time.
Agilegan: Stylizing portraits by inversion-consistent transfer learning
Guoxian Song, Linjie Luo, Jing Liu, Wan-Chun Ma, Chunpong Lai, Chuanxia Zheng, and Tat-Jen Cham · 2021
Closest in time.
Designing an encoder for stylegan image manipulation
Omer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen-Or · 2021
Closest in time.
On data augmentation for gan training
Ngoc-Trung Tran, Viet-Hung Tran, Ngoc-Bao Nguyen, Trung-Kien Nguyen, and Ngai-Man Cheung · 2021
Closest in time.
Regularizing generative adversarial networks under limited data
Hung-Yu Tseng, Lu Jiang, Ce Liu, Ming-Hsuan Yang, and Weilong Yang · 2021
Closest in time.
Cross-domain and disentangled face manipulation with 3d guidance, 2021
Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao · 2021
Closest in time.
Stylealign: Analysis and applications of aligned stylegan models, 2021
Zongze Wu, Yotam Nitzan, Eli Shechtman, and Dani Lischinski · 2021
Closest in time.
Gan inversion: A survey, 2021
Weihao Xia, Yulun Zhang, Yujiu Yang, Jing-Hao Xue, Bolei Zhou, and Ming-Hsuan Yang · 2021
Closest in time.
Generative hierarchical features from synthesizing images
Yinghao Xu, Yujun Shen, Jiapeng Zhu, Ceyuan Yang, and Bolei Zhou · 2021
Closest in time.
Data-efficient instance generation from instance discrimination
Ceyuan Yang, Yujun Shen, Yinghao Xu, and Bolei Zhou · 2021
Closest in time.
Gan prior embedded network for blind face restoration in the wild
Tao Yang, Peiran Ren, Xuansong Xie, and Lei Zhang · 2021
Closest in time.