Free-form image inpainting with gated convolution
Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang · 2019
Later among the works it cites.
Dm-gan: Dynamic memory generative adversarial networks for text-to-image synthesis
Minfeng Zhu, Pingbo Pan, Wei Chen, and Yi Yang · 2019
Later among the works it cites.
Language models are few-shot learners
Original
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever · 2020
Later among the works it cites.
X-lxmert: Paint, caption and answer questions with multi-modal transformers
Original
Jaemin Cho, Jiasen Lu, Dustin Schwenk, Hannaneh Hajishirzi, and Aniruddha Kembhavi · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Original
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Incorporating bert into parallel sequence decoding with adapters
Junliang Guo, Zhirui Zhang, Linli Xu, Hao-Ran Wei, Boxing Chen, and Enhong Chen · 2020
Later among the works it cites.
Non-autoregressive image captioning with counterfactuals-critical multi-agent learning
Original
Longteng Guo, Jing Liu, Xinxin Zhu, Xingjian He, Jie Jiang, and Hanqing Lu · 2020
Later among the works it cites.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Original
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu · 2020
Later among the works it cites.
Manigan: Text-guided image manipulation
Bowen Li, Xiaojuan Qi, Thomas Lukasiewicz, and Philip HS Torr · 2020
Later among the works it cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Ming Gong, and Daxin Jiang · 2020
Later among the works it cites.
Probabilistically masked language model capable of autoregressive generation in arbitrary word order
Original
Yi Liao, Xin Jiang, and Qun Liu · 2020
Later among the works it cites.
12-in-1: Multi-task vision and language representation learning
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee · 2020
Later among the works it cites.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Original
Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti · 2020
Later among the works it cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Original
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2020
Later among the works it cites.
Df-gan: Deep fusion generative adversarial networks for text-to-image synthesis
Original
Ming Tao, Hao Tang, Songsong Wu, Nicu Sebe, Xiao-Yuan Jing, Fei Wu, and Bingkun Bao · 2020
Later among the works it cites.
Text-guided image inpainting
Zijian Zhang, Zhou Zhao, Zhu Zhang, Baoxing Huai, and Jing Yuan · 2020
Later among the works it cites.
Unified vision-language pre-training for image captioning and vqa
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao · 2020
Later among the works it cites.
Sean: Image synthesis with semantic region-adaptive normalization
Peihao Zhu, Rameen Abdal, Yipeng Qin, and Peter Wonka · 2020
Later among the works it cites.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Björn Ommer · 2021
Closest in time.
M6: A chinese multimodal pretrainer
Original
Junyang Lin, Rui Men, An Yang, Chang Zhou, Ming Ding, Yichang Zhang, Peng Wang, Ang Wang, Le Jiang, Xianyan Jia, et al · 2021
Closest in time.
Zero-shot text-to-image generation
Original
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Closest in time.
Tedigan: Text-guided diverse image generation and manipulation
Weihao Xia, Yujiu Yang, Jing-Hao Xue, and Baoyuan Wu · 2021
Closest in time.
Exploring sparse expert models and beyond
Original
An Yang, Junyang Lin, Rui Men, Chang Zhou, Le Jiang, Xianyan Jia, Ang Wang, Jie Zhang, Jiamang Wang, Yong Li, et al · 2021
Closest in time.