Fetching the paper…
Reading the bibliography…
Multi-modal pretraining for learning high-level multi-modal representation is a further step towards deep learning and artificial intelligence.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
The Fifth PASCAL Recognizing Textual Entailment Challenge. In TAC 2009
Luisa Bentivogli, Bernardo Magnini, Ido Dagan, Hoa Trang Dang, and Danilo Giampiccolo. 2009 · 2009
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database. In CVPR 2009 . 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li. 2009 · 2009
Earlier work this paper cites.
Im2Text: Describing Images Using 1 Million Captioned Photographs. In NeurIPS 2011 . 1143–1151
Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. 2011 · 2011
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks. In NeurIPS 2012 . 1106–1114
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012 · 2012
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP 2013 . 1631–1642
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context. In ECCV 2014 . 740–755
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks. In NeurIPS 2014 . 3104–3112
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate. In ICLR 2015
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In NeurIPS 2015 . 91–99
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
A Neural Attention Model for Abstractive Sentence Summarization. In EMNLP 2015 . 379–389
Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition. In ICLR 2015
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In CVPR 2016 . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Bridging Nonlinearities and Stochastic Regularizers with Gaussian Error Linear Units
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, CoNLL 2016 . 280–290
Ramesh Nallapati, Bowen Zhou, Cícero Nogueira dos Santos, Çaglar Gülçehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100, 000+ Questions for Machine Comprehension of Text. In EMNLP 2016 . 2383–2392
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Learning Deep Structure-Preserving Image-Text Embeddings. In CVPR 2016 . 5005–5013
Liwei Wang, Yin Li, and Svetlana Lazebnik. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation. In SemEval@ACL 2017 . 1–14
Daniel M. Cer, Mona T. Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017 · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017 · 2017
Cited alongside, same era.
Attention is All you Need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017 . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In CVPR 2018 . 6077–6086
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Universal Language Model Fine-tuning for Text Classification. In ACL 2018 . 328–339
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Cited alongside, same era.
Stacked Cross Attention for Image-Text Matching. In ECCV 2018 . 212–228
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. 2018 · 2018
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. In NeurIPS 2019 . 13–23
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019a · 2019
Later among the works it cites.
12-in-1: Multi-Task Vision and Language Representation Learning
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee. 2019b · 2019
Later among the works it cites.
Natural Language Understanding with the Quora Question Pairs Dataset
Lakshay Sharma, Laura Graesser, Nikita Nangia, and Utku Evci. 2019 · 2019
Later among the works it cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2019 · 2019
Later among the works it cites.
Contrastive Bidirectional Transformer for Temporal Representation Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep Contextualized Word Representations. In NAACL-HLT 2018 . 2227–2237
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In ACL 2018 . 2556–2565
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Cited alongside, same era.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In NAACL-HLT 2018 . 1112–1122
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Cited alongside, same era.
Unified Vision-Language Pre-Training for Image Captioning and VQA. In AAAI 2020 . 13041–13049
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. 2020 · 2018
Cited alongside, same era.
Fusion of Detected Objects in Text for Visual Question Answering. In EMNLP-IJCNLP 2019 . 2131–2140
Chris Alberti, Jeffrey Ling, Michael Collins, and David Reitter. 2019 · 2019
Cited alongside, same era.
UNITER: Learning UNiversal Image-TExt Representations
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019 · 2019
Cited alongside, same era.
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid. 2019a · 2019
Later among the works it cites.
VideoBERT: A Joint Model for Video and Language Representation Learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019b · 2019
Later among the works it cites.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers. In EMNLP-IJCNLP 2019 . 5099–5110
Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In ICLR 2019
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Neural Network Acceptability Judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 2019
Later among the works it cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding. In NeurIPS 2019 . 5754–5764
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
From Recognition to Cognition: Visual Commonsense Reasoning. In CVPR 2019 . 6720–6731
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Later among the works it cites.
Large-Scale Adversarial Training for Vision-and-Language Representation Learning
Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. 2020 · 2020
Closest in time.
Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. 2020 · 2020
Closest in time.
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. 2020 · 2020
Closest in time.
End-to-End Learning of Visual Representations From Uncurated Instructional Videos. In CVPR 2020 . 9876–9886
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. 2020a · 2020
Closest in time.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti. 2020 · 2020
Closest in time.
Are we pretraining it right? Digging deeper into visio-linguistic pretraining
Amanpreet Singh, Vedanuj Goswami, and Devi Parikh. 2020 · 2020
Closest in time.
ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2020 · 2020
Closest in time.