Fetching the paper…
Reading the bibliography…
A method for creating a vision-and-language (V&L) model is to extend a language model through structural modifications and V&L pre-training.
Meetup! a corpus of joint activity dialogues in a visual environment
Nikolai Ilinykh, Sina Zarrieß, and David Schlangen. 2019 · 1907
Earlier work this paper cites.
VisualBERT: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 1908
Earlier work this paper cites.
The PhotoBook Dataset: Building Common Ground through Visually-Grounded Dialogue
Janosch Haber, Tim Baumgärtner, Ece Takmaz, Lieke Gelderloos, Elia Bruni, and Raquel Fernández. 2019 · 1910
Earlier work this paper cites.
Transformer is All You Need: Multimodal Multitask Learning with a Unified Transformer
Ronghang Hu and Amanpreet Singh. 2021 · 2003
Earlier work this paper cites.
M6-v0: Vision-and-Language Interaction for Multi-modal Pretraining
Junyang Lin, An Yang, Yichang Zhang, Jie Liu, Jingren Zhou, and Hongxia Yang. 2021 · 2003
Earlier work this paper cites.
The PASCAL Recognising Textual Entailment Challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
Automatically Constructing a Corpus of Sentential Paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
Roy Bar Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. 2006 · 2006
Earlier work this paper cites.
The Third PASCAL Recognizing Textual Entailment Challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007 · 2007
Earlier work this paper cites.
The Fifth PASCAL Recognizing Textual Entailment Challenge
Luisa Bentivogli, Ido Dagan, Hoa Trang Dang, Danilo Giampiccolo, and Bernardo Magnini. 2009 · 2009
Earlier work this paper cites.
Multimodal Pretraining Unmasked: Unifying the Vision and Language BERTs
Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, and Desmond Elliott. 2020 · 2011
Earlier work this paper cites.
The Winograd Schema Challenge
Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning
Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. 2020 · 2012
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
ReferItGame: Referring to Objects in Photographs of Natural Scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. 2014 · 2014
Cited alongside, same era.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Cited alongside, same era.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Like a baby: Visually situated neural language acquisition
Alexander Ororbia, Ankur Mali, Matthew Kelly, and David Reitter. 2019 · 2019
Later among the works it cites.
A Corpus for Reasoning about Natural Language Grounded in Photographs
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019 · 2019
Later among the works it cites.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
A natural language corpus of common grounding under continuous and partially-observable context
Takuma Udagawa and Akiko Aizawa. 2019 · 2019
Later among the works it cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Neural Network Acceptability Judgments
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi. 2017 · 2017
Cited alongside, same era.
The RepEval 2017 Shared Task: Multi-Genre Natural Language Inference with Sentence Representations
Nikita Nangia, Adina Williams, Angeliki Lazaridou, and Samuel Bowman. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Cited alongside, same era.
RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes
Semih Yagcioglu, Aykut Erdem, Erkut Erdem, and Nazli Ikizler-Cinbis. 2018 · 2018
Cited alongside, same era.
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Behind the scene: Revealing the secrets of pre-trained vision-and-language models
Jize Cao, Zhe Gan, Yu Cheng, Licheng Yu, Yen-Chun Chen, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
UNITER: Learning UNiversal image-TExt representations
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
ManyModalQA: Modality Disambiguation and QA over Diverse Inputs
Darryl Hannan, Akshay Jain, and Mohit Bansal. 2020 · 2020
Later among the works it cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-Art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
VisualMRC: Machine Reading Comprehension on Document Images
Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida. 2021 · 2021
Closest in time.