Fetching the paper…
Reading the bibliography…
Most existing research on visual question answering (VQA) is limited to information explicitly present in an image or a video.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 1908
Earlier work this paper cites.
Procedures as a representation for data in a computer program for understanding natural language
Terry Winograd. 1971 · 1971
Earlier work this paper cites.
Following instructions by imagining and reaching visual goals
John Kanu, Eadom Dessalene, Xiaomin Lin, Cornelia Fermuller, and Yiannis Aloimonos. 2020 · 2001
Earlier work this paper cites.
Learning what makes a difference from counterfactual examples and gradient supervision
Damien Teney, Ehsan Abbasnedjad, and Anton van den Hengel. 2020 · 2004
Earlier work this paper cites.
Graph edit distance reward: Learning to edit scene graph
Lichang Chen, Guosheng Lin, Shijie Wang, and Qingyao Wu. 2020 · 2008
Earlier work this paper cites.
Machine-learning maestro michael jordan on the delusions of big data and other huge engineering efforts
Lee Gomes. 2014 · 2014
Earlier work this paper cites.
Generative adversarial networks
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard Zemel. 2015 · 2015
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016 · 2016
Earlier work this paper cites.
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2017 · 2017
Earlier work this paper cites.
Semantic image synthesis via adversarial learning
Hao Dong, Simiao Yu, Chao Wu, and Yike Guo. 2017 · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. 2017 · 2017
Cited alongside, same era.
First quora dataset release: Question pairs
Shankar Iyer, Nikhil Dandekar, and Kornél Csernai. 2017 · 2017
Cited alongside, same era.
Figureqa: An annotated figure dataset for visual reasoning
Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Ákos Kádár, Adam Trischler, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi. 2017 · 2017
Cited alongside, same era.
Race: Large-scale reading comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A. Hudson and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
Clevr-dialog: A diagnostic dataset for multi-round reasoning in visual dialog
Satwik Kottur, José M. F. Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach. 2019 · 2019
Later among the works it cites.
Vision-based navigation with language-based assistance via imitation learning with indirect intervention
Khanh Nguyen, Debadeepta Dey, Chris Brockett, and Bill Dolan. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Mixture models for diverse machine translation: Tricks of the trade
Tianxiao Shen, Myle Ott, Michael Auli, and Marc’Aurelio Ranzato. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel. 2018 · 2018
Cited alongside, same era.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi. 2018 · 2018
Cited alongside, same era.
Iqa: Visual question answering in interactive environments
Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi. 2018 · 2018
Cited alongside, same era.
Dvqa: Understanding data visualizations via question answering
Kushal Kafle, Brian Price, Scott Cohen, and Christopher Kanan. 2018 · 2018
Cited alongside, same era.
Text-adaptive generative adversarial networks: Manipulating images with natural language
Seonghyeon Nam, Yunji Kim, and Seon Joo Kim. 2018 · 2018
Cited alongside, same era.
Answering visual what-if questions: From actions to predicted scene descriptions
Misha Wagner, Hector Basevi, Rakshith Shetty, Wenbin Li, Mateusz Malinowski, Mario Fritz, and Ales Leonardis. 2018 · 2018
Cited alongside, same era.
A dataset and architecture for visual reasoning with a working memory
Guangyu Robert Yang, Igor Ganichev, Xiao-Jing Wang, Jonathon Shlens, and David Sussillo. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Wiqa: A dataset for "what if…" reasoning over procedural text
Niket Tandon, Bhavana Dalvi Mishra, Keisuke Sakaguchi, Antoine Bosselut, and Peter Clark. 2019 · 2019
Later among the works it cites.
Composing text and image for image retrieval-an empirical odyssey
Nam Vo, Lu Jiang, Chen Sun, Kevin Murphy, Li-Jia Li, Li Fei-Fei, and James Hays. 2019 · 2019
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Later among the works it cites.
Video2commonsense: Generating commonsense descriptions to enrich video captioning
Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral, and Yezhou Yang. 2020 · 2020
Later among the works it cites.
Scene graph modification based on natural language commands
Xuanli He, Quan Hung Tran, Gholamreza Haffari, Walter Chang, Trung Bui, Zhe Lin, Franck Dernoncourt, and Nhan Dam. 2020 · 2020
Later among the works it cites.
What does BERT with vision look at?
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2020 · 2020
Later among the works it cites.
Visualcomet: Reasoning about the dynamic context of a still image
Jae Sung Park, Chandra Bhagavatula, Roozbeh Mottaghi, Ali Farhadi, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. 2020 · 2020
Later among the works it cites.