Fetching the paper…
Reading the bibliography…
We demonstrate the surprising strength of unimodal baselines in multimodal domains, and make concrete recommendations for best practices in future research.
Walk the talk: Connecting language, knowledge, and action in route instructions
Matt MacMahon, Brian Stankiewicz, and Benjamin Kuipers. 2006 · 2006
Earlier work this paper cites.
On comparing the power of mobile robots
Jason M O’Kane and Steven M. LaValle. 2006 · 2006
Earlier work this paper cites.
Learning to interpret natural language navigation instructions from observations
David L. Chen and Raymond J. Mooney. 2011 · 2011
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Exploring nearest neighbor approaches for image captioning
Jacob Devlin, Saurabh Gupta, Ross Girshick, Margaret Mitchell, and C Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
Hongyuan Mei, Mohit Bansal, and Matthew R. Walter. 2016 · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. 2016 · 2016
Earlier work this paper cites.
AI2-THOR: An Interactive 3D Environment for Visual AI
Eric Kolve, Roozbeh Mottaghi, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi. 2017 · 2017
Earlier work this paper cites.
Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel. 2018 · 2018
Earlier work this paper cites.
Mapping navigation instructions to continuous control actions with position visitation prediction
Valts Blukis, Dipendra Misra, Ross A. Knepper, and Yoav Artzi. 2018 · 2018
Earlier work this paper cites.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
Howard Chen, Alane Shur, Dipendra Misra, Noah Snavely, and Yoav Artzi. 2018 · 2018
Cited alongside, same era.
Visual referring expression recognition: What do systems actually learn?
Volkan Cirik, Louis-Philippe Morency, and Taylor Berg-Kirkpatrick. 2018 · 2018
Cited alongside, same era.
Speaker-follower models for vision-and-language navigation
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. 2018 · 2018
Cited alongside, same era.
Breaking NLI systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018 · 2018
Cited alongside, same era.
Iqa: Visual question answering in interactive environments
Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi. 2018 · 2018
Cited alongside, same era.
Vision-based navigation with language-based assistance via imitation learning with indirect intervention
Khanh Nguyen, Debadeepta Dey, Chris Brockett, and Bill Dolan. 2018 · 2018
Closest in time.
Hypothesis Only Baselines in Natural Language Inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Closest in time.
Object hallucination in image captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018 · 2018
Closest in time.
GQA: a new dataset for compositional question answering over real-world images
Drew A. Hudson and Christopher D. Manning. 2019 · 2019
Closest in time.
Tactical rewind: Self-correction via backtracking in vision-and-language navigation
Liyiming Ke, Xiujun Li, Yonatan Bisk, Ari Holtzman, Zhe Gan, Jingjing Liu, Jianfeng Gao, Yejin Choi, and Siddhartha Srinivasa. 2019 · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Aishwarya Agrawal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2018 · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
Divyansh Kaushik and Zachary C. Lipton. 2018 · 2018
Cited alongside, same era.
Did the model understand the question?
Pramod Kaushik Mudrakarta, Ankur Taly, Mukund Sundararajan, and Kedar Dhamdhere. 2018 · 2018
Cited alongside, same era.
Embodied Question Answering
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra. 2018a
Cited in the paper.
Neural Modular Control for Embodied Question Answering
Abhishek Das, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra. 2018b
Cited in the paper.
Self-aware visual-textual co-grounded navigation agent
Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan AlRegib, Zsolt Kira, Richard Socher, and Caiming Xiong. 2019a
Cited in the paper.
Chris Paxton, Yonatan Bisk, Jesse Thomason, Arunkumar Byravan, and Dieter Fox. 2019 · 2019
Closest in time.
Learning to navigate unseen environments: Back translation with environmental dropout
Hao Tan, Licheng Yu, and Mohit Bansal. 2019 · 2019
Closest in time.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. 2019 · 2019
Closest in time.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Closest in time.