Fetching the paper…
Reading the bibliography…
In this work, we construct the largest dataset for multimodal pretraining in Chinese, which consists of over 1.9TB images and 292GB texts that cover a wide range of domains.
Taming Transformers for High-Resolution Image Synthesis
Patrick Esser, Robin Rombach, and Björn Ommer. 2020 · 2012
Earlier work this paper cites.
Generative adversarial networks
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context. In ECCV 2014 . 740–755
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question answering
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In NeurIPS 2015 . 91–99
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In CVPR 2016 . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units. In ACL 2016
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Thuctc: an efficient chinese text classifier
Maosong Sun, Jingyang Li, Zhipeng Guo, Z Yu, Y Zheng, X Si, and Z Liu. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning. In NIPS
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In NeurIPS 2017 . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks. In CVPR 2017 . 1492–1500
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. 2017 · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Earlier work this paper cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1316–1324
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT 2019 . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Unified Language Model Pre-training for Natural Language Understanding and Generation. In NeurIPS 2019 . 13042–13054
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 2019
Earlier work this paper cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Cited alongside, same era.
Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training
Gen Li, Nan Duan, Yuejian Fang, Daxin Jiang, and Ming Zhou. 2019a · 2019
Cited alongside, same era.
VisualBERT: A Simple and Performant Baseline for Vision and Language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019b · 2019
Cited alongside, same era.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Large-Scale Adversarial Training for Vision-and-Language Representation Learning. In NeurIPS 2020
Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Fashionbert: Text and image matching with adaptive loss for cross-modal retrieval. In SIGIR 2020 . 2251–2260
Dehong Gao, Linbo Jin, Ben Chen, Minghui Qiu, Peng Li, Yi Wei, Yi Hu, and Hao Wang. 2020 · 2020
Later among the works it cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020 · 2020
Later among the works it cites.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. 2020 · 2020
Later among the works it cites.
Convbert: Improving bert with span-based dynamic convolution
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. In NeurIPS 2019 . 13–23
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019a · 2019
Cited alongside, same era.
12-in-1: Multi-Task Vision and Language Representation Learning
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee. 2019b · 2019
Cited alongside, same era.
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 2019
Cited alongside, same era.
MASS: Masked Sequence to Sequence Pre-training for Language Generation. In ICML 2019 . 5926–5936
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019 · 2019
Cited alongside, same era.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers. In EMNLP-IJCNLP 2019 . 5099–5110
Hao Tan and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
Structbert: Incorporating language structures into pre-training for deep language understanding
Wei Wang, Bin Bi, Ming Yan, Chen Wu, Zuyi Bao, Jiangnan Xia, Liwei Peng, and Luo Si. 2019 · 2019
Cited alongside, same era.
XLNet: Generalized Autoregressive Pretraining for Language Understanding. In NeurIPS 2019 . 5754–5764
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Cited alongside, same era.
Unilmv2: Pseudo-masked language models for unified language model pre-training. In International Conference on Machine Learning . PMLR, 642–652
Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Jianfeng Gao, Songhao Piao, Ming Zhou, et al · 2020
Cited alongside, same era.
Zihang Jiang, Weihao Yu, Daquan Zhou, Yunpeng Chen, Jiashi Feng, and Shuicheng Yan. 2020 · 2020
Later among the works it cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020 · 2020
Later among the works it cites.
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
Interbert: Vision-and-language interaction for multi-modal pretraining
Junyang Lin, An Yang, Yichang Zhang, Jie Liu, Jingren Zhou, and Hongxia Yang. 2020 · 2020
Later among the works it cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations. In ICLR 2020
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
Whale: A Unified Distributed Training Framework
Ang Wang, Xianyan Jia, Le Jiang, Jie Zhang, Yong Li, and Wei Lin. 2020 · 2020
Later among the works it cites.
CLUECorpus2020: A Large-scale Chinese Corpus for Pre-trainingLanguage Model
Liang Xu, Xuanwei Zhang, and Qianqian Dong. 2020 · 2020
Later among the works it cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2020 · 2020
Later among the works it cites.
Unified Vision-Language Pre-Training for Image Captioning and VQA. In AAAI 2020 . 13041–13049
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2021 · 2021
Closest in time.
BASE Layers: Simplifying Training of Large, Sparse Models
Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer. 2021 · 2021
Closest in time.
Zero-Shot Text-to-Image Generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Closest in time.