Fetching the paper…
Reading the bibliography…
We review research on generating visual data from text from the angle of "cross-modal generation." This point of view allows us to draw parallels between various methods geared towards working on input text and producing visual output, without limiting the analysis to narrow sub-areas.
Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 1877–1901
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Generating Long Sequences with Sparse Transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation
Lincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding, Yixing Zheng, Xin Yu, and Changjie Fan. 2021 · 1920
Earlier work this paper cites.
The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. 2020 · 1981
Earlier work this paper cites.
A learning algorithm for boltzmann machines
David H. Ackley, Geoffrey E. Hinton, and Terrence J. Sejnowski. 1985 · 1985
Earlier work this paper cites.
Signature Verification using a "Siamese" Time Delay Neural Network. In Advances in Neural Information Processing Systems , J. Cowan, G. Tesauro, and J. Alspector (Eds.), Vol. 6. Morgan-Kaufmann
Jane Bromley, Isabelle Guyon, Yann LeCun, Eduard Säckinger, and Roopak Shah. 1993 · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. 1998 · 1998
Earlier work this paper cites.
Training Products of Experts by Minimizing Contrastive Divergence
Geoffrey E. Hinton. 2002 · 2002
Earlier work this paper cites.
Recognizing human actions: a local SVM approach. In Proceedings of the 17th International Conference on Pattern Recognition (ICPR) , Vol. 3. 32–36
C. Schuldt, I. Laptev, and B. Caputo. 2004 · 2004
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05) , Vol. 1. 539–546
S. Chopra, R. Hadsell, and Y. LeCun. 2005 · 2005
Earlier work this paper cites.
Labeled Faces in the Wild: A Database forStudying Face Recognition in Unconstrained Environments. In Workshop on Faces in ’Real-Life’ Images: Detection, Alignment, and Recognition . Erik Learned-Miller and Andras Ferencz and Frédéric Jurie, Marseille, France
Gary B. Huang, Marwan Mattar, Tamara Berg, and Eric Learned-Miller. 2008 · 2008
Earlier work this paper cites.
Introduction to Information Retrieval
Christopher D. Manning, Prabhakar Raghavan, and Hinrich Schütze. 2008 · 2008
Earlier work this paper cites.
Automated Flower Classification over a Large Number of Classes. In Indian Conference on Computer Vision, Graphics and Image Processing
Maria-Elena Nilsback and Andrew Zisserman. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on Computer Vision and Pattern Recognition . 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
The MUG facial expression database. In 11th International Workshop on Image Analysis for Multimedia Interactive Services WIAMIS 10 . 1–4
Niki Aifanti, Christos Papachristou, and Anastasios Delopoulos. 2010 · 2010
Earlier work this paper cites.
Reading Digits in Natural Images with Unsupervised Feature Learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng. 2011 · 2011
Earlier work this paper cites.
Technical Report CNS-TR-2011-001
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. 2011 · 2011
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks. In Proceedings of the 25th International Conference on Neural Information Processing Systems - Volume 1 (Lake Tahoe, Nevada) (NIPS’12) . Curran Associates Inc., Red Hook, NY, USA, 1097–1105
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012 · 2012
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Association for Computational Linguistics, Brussels, Belgium, 66–71
Taku Kudo and John Richardson. 2018 · 2012
Earlier work this paper cites.
Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville. 2013 · 2013
Earlier work this paper cites.
Bringing Semantics into Focus Using Visual Abstraction. In 2013 IEEE Conference on Computer Vision and Pattern Recognition . 3009–3016
C. Lawrence Zitnick and Devi Parikh. 2013 · 2013
Earlier work this paper cites.
Generative Adversarial Nets. In Advances in Neural Information Processing Systems , Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K. Q. Weinberger (Eds.), Vol. 27. Curran Associates, Inc
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Auto-Encoding Variational Bayes. In 2nd International Conference on Learning Representations, ICLR 2014
Diederik P. Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context. In Computer Vision – ECCV 2014 , David Fleet, Tomas Pajdla, Bernt Schiele, and Tinne Tuytelaars (Eds.). Springer International Publishing, 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings , Yoshua Bengio and Yann LeCun (Eds.)
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Microsoft COCO Captions: Data Collection and Evaluation Server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Convolutional Networks on Graphs for Learning Molecular Fingerprints. In Advances in Neural Information Processing Systems , C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett (Eds.), Vol. 28. Curran Associates, Inc
David K Duvenaud, Dougal Maclaurin, Jorge Iparraguirre, Rafael Bombarell, Timothy Hirzel, Alan Aspuru-Guzik, and Ryan P Adams. 2015 · 2015
Earlier work this paper cites.
Deep Learning Face Attributes in the Wild. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2015 · 2015
Earlier work this paper cites.
Deep Unsupervised Learning using Nonequilibrium Thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 37) , Francis Bach and David Blei (Eds.). PMLR, Lille, France, 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Unsupervised Learning of Video Representations Using LSTMs. In Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37 (Lille, France) (ICML’15) . JMLR.org, 843–852
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhutdinov. 2015 · 2015
Earlier work this paper cites.
A guide to convolution arithmetic for deep learning
Vincent Dumoulin and Francesco Visin. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Generating Images from Captions with Attention. In ICLR
Elman Mansimov, Emilio Parisotto, Jimmy Ba, and Ruslan Salakhutdinov. 2016 · 2016
Earlier work this paper cites.
Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks. In 4th International Conference on Learning Representations, ICLR 2016
Alec Radford, Luke Metz, and Soumith Chintala. 2016 · 2016
Earlier work this paper cites.
Improved Techniques for Training GANs. In Advances in Neural Information Processing Systems , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Eds.), Vol. 29. Curran Associates, Inc
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, Xi Chen, and Xi Chen. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, Berlin, Germany, 1715–1725
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Pixel Recurrent Neural Networks. In Proceedings of The 33rd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 48) , Maria Florina Balcan and Kilian Q. Weinberger (Eds.). PMLR, New York, New York, USA, 1747–1756
Aäron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Attribute2Image: Conditional Image Generation from Visual Attributes. In Computer Vision – ECCV 2016 , Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling (Eds.). Springer International Publishing, Cham, 776–791
Xinchen Yan, Jimei Yang, Kihyuk Sohn, and Honglak Lee. 2016 · 2016
Earlier work this paper cites.
LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop
Fisher Yu, Ari Seff, Yinda Zhang, Shuran Song, Thomas Funkhouser, and Jianxiong Xiao. 2016 · 2016
Earlier work this paper cites.
Show, Adapt and Tell: Adversarial Training of Cross-Domain Image Captioner. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
Tseng-Hung Chen, Yuan-Hong Liao, Ching-Yao Chuang, Wan-Ting Hsu, Jianlong Fu, and Min Sun. 2017 · 2017
Earlier work this paper cites.
Semantic Image Synthesis via Adversarial Learning. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
Hao Dong, Simiao Yu, Chao Wu, and Yike Guo. 2017 · 2017
Earlier work this paper cites.
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17) . Curran Associates Inc., Red Hook, NY, USA, 6629–6640
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017 · 2017
Earlier work this paper cites.
Semi-Supervised Classification with Graph Convolutional Networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings
Thomas N. Kipf and Max Welling. 2017 · 2017
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017 · 2017
Earlier work this paper cites.
Attentive Semantic Video Generation Using Captions. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
Tanya Marwah, Gaurav Mittal, and Vineeth N. Balasubramanian. 2017 · 2017
Earlier work this paper cites.
Sync-DRAW: Automatic Video Generation Using Deep Recurrent Attentive Architectures. In Proceedings of the 25th ACM International Conference on Multimedia (Mountain View, California, USA) (MM ’17) . Association for Computing Machinery, New York, NY, USA, 1096–1104
Gaurav Mittal, Tanya Marwah, and Vineeth N. Balasubramanian. 2017 · 2017
Earlier work this paper cites.
Plug & Play Generative Networks: Conditional Iterative Generation of Images in Latent Space. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Anh Nguyen, Jeff Clune, Yoshua Bengio, Alexey Dosovitskiy, and Jason Yosinski. 2017 · 2017
Earlier work this paper cites.
Joint Multimodal Learning with Deep Generative Models. In ICLR
Masahiro Suzuki, Kotaro Nakayama, and Yutaka Matsuo. 2017 · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17) . Curran Associates Inc., Red Hook, NY, USA, 6309–6318
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In Advances in Neural Information Processing Systems , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
StackGAN: Text to Photo-Realistic Image Synthesis With Stacked Generative Adversarial Networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N. Metaxas. 2017 · 2017
Earlier work this paper cites.
Semi-supervised FusedGAN for Conditional Image Generation. In Proceedings of the European Conference on Computer Vision (ECCV)
Navaneeth Bodla, Gang Hua, and Rama Chellappa. 2018 · 2018
Cited alongside, same era.
StarGAN: Unified Generative Adversarial Networks for Multi-Domain Image-to-Image Translation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. 2018 · 2018
Cited alongside, same era.
Image Generation from Scene Graphs. In CVPR
Justin Johnson, Agrim Gupta, and Li Fei-Fei. 2018 · 2018
Cited alongside, same era.
Video Generation From Text
Yitong Li, Martin Min, Dinghan Shen, David Carlson, and Lawrence Carin. 2018 · 2018
Cited alongside, same era.
DA-GAN: Instance-Level Image Translation by Deep Attention Generative Adversarial Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Shuang Ma, Jianlong Fu, Chang Wen Chen, and Tao Mei. 2018 · 2018
Language-Guided Global Image Editing via Cross-Modal Cyclic Mechanism. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 2115–2124
Wentao Jiang, Ning Xu, Jiayun Wang, Chen Gao, Jing Shi, Zhe Lin, and Si Liu. 2021 · 2021
Later among the works it cites.
Learning to Compose Visual Relations
Nan Liu, Shuang Li, Yilun Du, Josh Tenenbaum, and Antonio Torralba. 2021 · 2021
Later among the works it cites.
Improved Denoising Diffusion Probabilistic Models. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, 8162–8171
Alexander Quinn Nichol and Prafulla Dhariwal. 2021 · 2021
Later among the works it cites.
Controllable and compositional generation with latent-space energy-based models. In Thirty-Fifth Conference on Neural Information Processing Systems
Weili Nie, Arash Vahdat, and Anima Anandkumar. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Text-Adaptive Generative Adversarial Networks: Manipulating Images with Natural Language. In Advances in Neural Information Processing Systems , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.), Vol. 31. Curran Associates, Inc
Seonghyeon Nam, Yunji Kim, and Seon Joo Kim. 2018 · 2018
Cited alongside, same era.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In ACL . 2556–2565
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Cited alongside, same era.
MoCoGAN: Decomposing motion and content for video generation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 1526–1535
Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz. 2018 · 2018
Cited alongside, same era.
Multimodal Generative Models for Scalable Weakly-Supervised Learning. In Proceedings of the 32nd International Conference on Neural Information Processing Systems (Montréal, Canada) (NIPS’18) . Curran Associates Inc., Red Hook, NY, USA, 5580–5590
Mike Wu and Noah Goodman. 2018 · 2018
Cited alongside, same era.
AttnGAN: Fine-Grained Text to Image Generation With Attentional Generative Adversarial Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. 2018 · 2018
Cited alongside, same era.
StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial Networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N. Metaxas. 2019 · 2018
Cited alongside, same era.
Specifying Object Attributes and Relations in Interactive Scene Generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
Oron Ashual and Lior Wolf. 2019 · 2019
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Zero-Shot Text-to-Image Generation. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, 8821–8831
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Pivotal Tuning for Latent-based Editing of Real Images
Daniel Roich, Ron Mokady, Amit H Bermano, and Daniel Cohen-Or. 2021 · 2021
Later among the works it cites.
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs. In Advances in Neural Information Processing Systems
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021 · 2021
Later among the works it cites.
Denoising Diffusion Implicit Models. In International Conference on Learning Representations
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2021 · 2021
Later among the works it cites.
Generalized Multimodal ELBO. In International Conference on Learning Representations
Thomas M. Sutter, Imant Daunhawer, and Julia E Vogt. 2021 · 2021
Later among the works it cites.
MultiModalQA: complex question answering over text, tables and images. In International Conference on Learning Representations
Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant. 2021 · 2021
Later among the works it cites.
Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps Grounding. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 2720–2728
Dexin Wang and Deyi Xiong. 2021 · 2021
Later among the works it cites.
TediGAN: Text-Guided Diverse Face Image Generation and Manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 2256–2265
Weihao Xia, Yujiu Yang, Jing-Hao Xue, and Baoyuan Wu. 2021 · 2021
Later among the works it cites.
UniMF: A Unified Framework to Incorporate Multimodal Knowledge Bases into End-to-End Task-Oriented Dialogue Systems. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21 , Zhi-Hua Zhou (Ed.). International Joint Conferences on Artificial Intelligence Organization, 3978–3984
Shiquan Yang, Rui Zhang, Sarah M. Erfani, and Jey Han Lau. 2021b · 2021
Later among the works it cites.
Barlow Twins: Self-Supervised Learning via Redundancy Reduction. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, 12310–12320
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stephane Deny. 2021 · 2021
Later among the works it cites.
CrossCLR: Cross-Modal Contrastive Learning for Multi-Modal Video Representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 1450–1459
Mohammadreza Zolfaghari, Yi Zhu, Peter Gehler, and Thomas Brox. 2021 · 2021
Later among the works it cites.
Blended Diffusion for Text-Driven Editing of Natural Images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 18208–18218
Omri Avrahami, Dani Lischinski, and Ohad Fried. 2022 · 2022
Later among the works it cites.
Text2live: Text-driven layered image and video editing. In European Conference on Computer Vision . Springer, 707–723
Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, and Tali Dekel. 2022 · 2022
Later among the works it cites.
VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning. In International Conference on Learning Representations
Adrien Bardes, Jean Ponce, and Yann LeCun. 2022 · 2022
Later among the works it cites.
Retrieval-Augmented Diffusion Models. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 15309–15324
Andreas Blattmann, Robin Rombach, Kaan Oktay, Jonas Müller, and Björn Ommer. 2022 · 2022
Later among the works it cites.
Multimodal Adversarially Learned Inference with Factorized Discriminators. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 36. 6304–6312
Wenxue Chen and Jianke Zhu. 2022 · 2022
Later among the works it cites.
FlexIT: Towards Flexible Semantic Image Translation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Guillaume Couairon, Asya Grechka, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. 2022 · 2022
Later among the works it cites.
On the Limitations of Multimodal VAEs. In International Conference on Learning Representations
Imant Daunhawer, Thomas M. Sutter, Kieran Chin-Cheong, Emanuele Palumbo, and Julia E Vogt. 2022 · 2022
Later among the works it cites.
M3L: Language-Based Video Editing via Multi-Modal Multi-Level Transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 10513–10522
Tsu-Jui Fu, Xin Eric Wang, Scott T. Grafton, Miguel P. Eckstein, and William Yang Wang. 2022 · 2022
Later among the works it cites.
StyleGAN-NADA: CLIP-Guided Domain Adaptation of Image Generators. In SIGGRAPH
Rinon Gal, Or Patashnik, Haggai Maron, Gal Chechik, and Daniel Cohen-Or. 2022 · 2022
Later among the works it cites.
Vector Quantized Diffusion Model for Text-to-Image Synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 10696–10706
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. 2022 · 2022
Later among the works it cites.
Show Me What and Tell Me How: Video Synthesis via Multimodal Conditioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 3615–3625
Ligong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri, Kyle Olszewski, Shervin Minaee, Dimitris Metaxas, and Sergey Tulyakov. 2022 · 2022
Later among the works it cites.
MUGEN: A Playground for Video-Audio-Text Multimodal Understanding and GENeration. In Computer Vision - ECCV 2022 , Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner (Eds.). Springer Nature Switzerland, Cham, 431–449
Thomas Hayes, Songyang Zhang, Xi Yin, Guan Pang, Sasha Sheng, Harry Yang, Songwei Ge, Qiyuan Hu, and Devi Parikh. 2022 · 2022
Later among the works it cites.
Video Diffusion Models. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 8633–8646
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. 2022 · 2022
Later among the works it cites.
Multimodal Conditional Image Synthesis with Product-of-Experts GANs. In Computer Vision - ECCV 2022 , Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner (Eds.). Springer Nature Switzerland, Cham, 91–109
Xun Huang, Arun Mallya, Ting-Chun Wang, and Ming-Yu Liu. 2022 · 2022
Later among the works it cites.
Learning Multimodal VAEs through Mutual Supervision. In International Conference on Learning Representations
Tom Joy, Yuge Shi, Philip Torr, Tom Rainforth, Sebastian M Schmon, and Siddharth N. 2022 · 2022
Later among the works it cites.
Cross-modal Representation Learning and Relation Reasoning for Bidirectional Adaptive Manipulation. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22 , Lud De Raedt (Ed.). International Joint Conferences on Artificial Intelligence Organization, 3222–3228
Lei Li, Kai Fan, and Chun Yuan. 2022 · 2022
Later among the works it cites.
BEAT: A Large-Scale Semantic and Emotional Multi-modal Dataset for Conversational Gestures Synthesis. In Computer Vision - ECCV 2022 , Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner (Eds.). Springer Nature Switzerland, Cham, 612–630
Haiyang Liu, Zihao Zhu, Naoya Iwamoto, Yichen Peng, Zhengqing Li, You Zhou, Elif Bozkurt, and Bo Zheng. 2022b · 2022
Later among the works it cites.
Compositional Visual Generation with Composable Diffusion Models. In Computer Vision - ECCV 2022 , Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner (Eds.). Springer Nature Switzerland, Cham, 423–439
Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B. Tenenbaum. 2022a · 2022
Later among the works it cites.
Deep Learning Methods for Abstract Visual Reasoning: A Survey on Raven’s Progressive Matrices
Mikołaj Małkiński and Jacek Mańdziuk. 2022 · 2022
Later among the works it cites.
A review of emerging research directions in Abstract Visual Reasoning
Mikołaj Małkiǹski and Jacek Maǹdziuk. 2023 · 2022
Later among the works it cites.
SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations. In International Conference on Learning Representations
Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. 2022 · 2022
Later among the works it cites.
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. In Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 162) , Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (Eds.). PMLR, 16784–16804
Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob Mcgrew, Ilya Sutskever, and Mark Chen. 2022 · 2022
Later among the works it cites.
Hierarchical Text-Conditional Image Generation with CLIP Latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Later among the works it cites.
High-Resolution Image Synthesis With Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Later among the works it cites.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 36479–36494
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. 2022 · 2022
Later among the works it cites.
SemanticStyleGAN: Learning Compositional Generative Priors for Controllable Image Synthesis and Editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 11254–11264
Yichun Shi, Xiao Yang, Yangyue Wan, and Xiaohui Shen. 2022 · 2022
Later among the works it cites.
DF-GAN: A Simple and Effective Baseline for Text-to-Image Synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 16515–16525
Ming Tao, Hao Tang, Fei Wu, Xiao-Yuan Jing, Bing-Kun Bao, and Changsheng Xu. 2022 · 2022
Later among the works it cites.
FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Yan Wang, Yixuan Sun, Yiwen Huang, Zhongying Liu, Shuyong Gao, Wei Zhang, Weifeng Ge, and Wenqiang Zhang. 2022 · 2022
Later among the works it cites.
Generative Visual Prompt: Unifying Distributional Control of Pre-Trained Generative Models. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 22422–22437
Chen Henry Wu, Saman Motamed, Shaunak Srivastava, and Fernando D De la Torre. 2022 · 2022
Later among the works it cites.
Towards Language-Free Training for Text-to-Image Generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 17907–17917
Yufan Zhou, Ruiyi Zhang, Changyou Chen, Chunyuan Li, Chris Tensmeyer, Tong Yu, Jiuxiang Gu, Jinhui Xu, and Tong Sun. 2022 · 2022
Later among the works it cites.
InstructPix2Pix: Learning To Follow Image Editing Instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 18392–18402
Tim Brooks, Aleksander Holynski, and Alexei A. Efros. 2023 · 2023
Later among the works it cites.
Prompt-to-Prompt Image Editing with Cross-Attention Control. In The Eleventh International Conference on Learning Representations
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-or. 2023 · 2023
Later among the works it cites.
Imagic: Text-Based Real Image Editing With Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 6007–6017
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. 2023 · 2023
Later among the works it cites.
NULL-Text Inversion for Editing Real Images Using Guided Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 6038–6047
Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. 2023 · 2023
Later among the works it cites.
Dual Diffusion Implicit Bridges for Image-to-Image Translation. In International Conference on Learning Representations
Xuan Su, Jiaming Song, Chenlin Meng, and Stefano Ermon. 2023 · 2023
Later among the works it cites.
Cross-modal text and visual generation: A systematic review. Part 1: Image to text
Maciej Żelaszczyk and Jacek Mańdziuk. 2023 · 2023
Later among the works it cites.
StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 2085–2094
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. 2021 · 2094
Closest in time.