Fetching the paper…
Reading the bibliography…
Creation of images using generative adversarial networks has been widely adapted into multi-modal regime with the advent of multi-modal representation models pre-trained on large corpus.
Beat tracking by dynamic programming
Daniel PW Ellis · 2007
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Combining language and vision with a multimodal skip-gram model, 2015
Angeliki Lazaridou, Nghia The Pham, and Marco Baroni · 2015
Earlier work this paper cites.
Neural discrete representation learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Large scale gan training for high fidelity natural image synthesis, 2019
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks, 2019
Tero Karras, Samuli Laine, and Timo Aila · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks, 2019
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Deep music visualizer
msieg · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers, 2019
Hao Tan and Mohit Bansal · 2019
Earlier work this paper cites.
Uniter: Universal image-text representation learning, 2020
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Editing in style: Uncovering the local semantics of gans
Edo Collins, Raja Bala, Bob Price, and Sabine Susstrunk · 2020
Cited alongside, same era.
Taming transformers for high-resolution image synthesis, 2020
Patrick Esser, Robin Rombach, and Björn Ommer · 2020
Cited alongside, same era.
Ganspace: Discovering interpretable gan controls
Erik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, and Sylvain Paris · 2020
Cited alongside, same era.
On the ”steerability” of generative adversarial networks
Ali Jahanian*, Lucy Chai*, and Phillip Isola · 2020
Cited alongside, same era.
Analyzing and improving the image quality of stylegan, 2020
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila · 2020
Cogview: Mastering text-to-image generation via transformers, 2021
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang · 2021
Later among the works it cites.
Fusedream: Training-free text-to-image generation with improved clip+gan space optimization, 2021
Xingchao Liu, Chengyue Gong, Lemeng Wu, Shujian Zhang, Hao Su, and Qiang Liu · 2021
Later among the works it cites.
Vqgan-clip
nerdyrodent · 2021
Later among the works it cites.
Styleclip: Text-driven manipulation of stylegan imagery
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Interpreting the latent space of gans for semantic face editing
Yujun Shen, Jinjin Gu, Xiaoou Tang, and Bolei Zhou · 2020
Cited alongside, same era.
Unsupervised discovery of interpretable directions in the gan latent space
Andrey Voynov and Artem Babenko · 2020
Cited alongside, same era.
Styleflow: Attribute-conditioned exploration of stylegan-generated images using conditional continuous normalizing flows
Rameen Abdal, Peihao Zhu, Niloy J Mitra, and Peter Wonka · 2021
Cited alongside, same era.
Yujun Shen and Bolei Zhou · 2021
Later among the works it cites.
Clip music video
VQGAN+CLIP · 2021
Later among the works it cites.
A geometric analysis of deep generative image models and its applications
Binxu Wang and Carlos R Ponce · 2021
Later among the works it cites.
Wav2clip: Learning robust audio representations from clip, 2021
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello · 2021
Later among the works it cites.