Fetching the paper…
Reading the bibliography…
We introduce Room-Across-Room (RxR), a new Vision-and-Language Navigation (VLN) dataset.
VALAN: Vision and language agent navigation
Larry Lansing, Vihan Jain, Harsh Mehta, Haoshuo Huang, and Eugene Ie. 2019 · 1912
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Spatial language and spatial representation: A cross-linguistic comparison
Edward Munnich, Barbara Landau, and Barbara Anne Dosher. 2001 · 2001
Earlier work this paper cites.
Crisscrossed captions: Extended intramodal and intermodal semantic similarity judgments for MS-COCO
Zarana Parekh, Jason Baldridge, Daniel Cer, Austin Waters, and Yinfei Yang. 2020 · 2004
Earlier work this paper cites.
Massively multilingual asr: 50 languages, 1 model, 1 billion parameters
Vineel Pratap, Anuroop Sriram, Paden Tomasello, Awni Hannun, Vitaliy Liptchinsky, Gabriel Synnaeve, and Ronan Collobert. 2020 · 2007
Earlier work this paper cites.
Linguistically naïve != language independent: Why NLP needs linguistic typology
Emily M. Bender. 2009 · 2009
Earlier work this paper cites.
Learning to interpret natural language navigation instructions from observations
David L Chen and Raymond J Mooney. 2011 · 2011
Earlier work this paper cites.
Plasticity of human spatial cognition: Spatial language and cognition covary across cultures
Daniel BM Haun, Christian J Rapold, Gabriele Janzen, and Stephen C Levinson. 2011 · 2011
Earlier work this paper cites.
Mapping spatial frames of reference onto time: A review of theoretical accounts and empirical findings
Andrea Bender and Sieghard Beller. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Matterport3D: Learning from RGB-D data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. 2017 · 2017
Cited alongside, same era.
Speaker-follower models for vision-and-language navigation
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. 2018 · 2018
Cited alongside, same era.
Mapping instructions to actions in 3d environments with visual goal prediction
Dipendra Kumar Misra, Andrew Bennett, Valts Blukis, Eyvind Niklasson, Max Shatkhin, and Yoav Artzi. 2018 · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Cited alongside, same era.
Massively multilingual neural machine translation
Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019 · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le. 2019 · 2019
Later among the works it cites.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. 2019 · 2019
Later among the works it cites.
Self-critical reasoning for robust visual question answering
Jialin Wu and Raymond Mooney. 2019 · 2019
Later among the works it cites.
TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020 · 2020
Closest in time.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumeé III, and Kate Crawford. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Touchdown: Natural language navigation and spatial reasoning in visual street environments
Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Transferable representation learning in vision-and-language navigation
Haoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku, Gabriel Magalhaes, Jason Baldridge, and Eugene Ie. 2019 · 2019
Cited alongside, same era.
Effective and general evaluation for instruction conditioned navigation using dynamic time warping
Gabriel Ilharco, Vihan Jain, Alexander Ku, Eugene Ie, and Jason Baldridge. 2019 · 2019
Cited alongside, same era.
Stay on the path: Instruction fidelity in vision-and-language navigation
Vihan Jain, Gabriel Magalhaes, Alexander Ku, Ashish Vaswani, Eugene Ie, and Jason Baldridge. 2019 · 2019
Cited alongside, same era.
Robust navigation with language pretraining and stochastic sampling
Xiujun Li, Chunyuan Li, Qiaolin Xia, Yonatan Bisk, Asli Celikyilmaz, Jianfeng Gao, Noah Smith, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Taking a hint: Leveraging explanations to make vision and language models more grounded
Ramprasaath R Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Shalini Ghosh, Larry Heck, Dhruv Batra, and Devi Parikh. 2019 · 2019
Cited alongside, same era.
Sub-instruction aware vision-and-language navigation
Yicong Hong, Cristian Rodriguez-Opazo, Qi Wu, and Stephen Gould. 2020 · 2020
Closest in time.
Beyond the nav-graph: Vision and language navigation in continuous environments
Jacob Krantz, Erik Wijmans, Arjun Majundar, Dhruv Batra, and Stefan Lee. 2020 · 2020
Closest in time.
Retouchdown: Adding touchdown to streetlearn as a shareable resource for language grounding tasks in street view
Harsh Mehta, Yoav Artzi, Jason Baldridge, Eugene Ie, and Piotr Mirowski. 2020 · 2020
Closest in time.
Connecting vision and language with localized narratives
Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut, and Vittorio Ferrari. 2020 · 2020
Closest in time.
REVERIE: Remote embodied visual referring expression in real indoor environments
Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. 2020 · 2020
Closest in time.
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. 2020 · 2020
Closest in time.
Environment-agnostic multitask learning for natural language grounded navigation
Xin Wang, Vihan Jain, Eugene Ie, William Yang Wang, Zornitsa Kozareva, and Sujith Ravi. 2020 · 2020
Closest in time.
BabyWalk: Going farther in vision-and-language navigation by taking baby steps
Wang Zhu, Hexiang Hu, Jiacheng Chen, Zhiwei Deng, Vihan Jain, Eugene Ie, and Fei Sha. 2020 · 2020
Closest in time.