Fetching the paper…
Reading the bibliography…
Embodied instruction following is a challenging problem requiring an agent to infer a sequence of primitive actions to achieve a goal environment state from complex language and visual inputs.
Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
Ma, C.-Y.; Lu, J.; Wu, Z.; AlRegib, G.; Kira, Z.; Socher, R.; and Xiong, C. 2019a · 1901
Earlier work this paper cites.
The Regretful Agent: Heuristic-Aided Navigation through Progress Estimation
Ma, C.-Y.; Wu, Z.; AlRegib, G.; Xiong, C.; and Kira, Z. 2019b · 1903
Earlier work this paper cites.
The Development of Spatial Representations of Large-Scale Environments
Siegel, A. W.; and White, S. H. 1975 · 1975
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
YOLOv4: Optimal Speed and Accuracy of Object Detection
Bochkovskiy, A.; Wang, C.-Y.; and Liao, H.-Y. M. 2020 · 2004
Earlier work this paper cites.
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Li, X.; Yin, X.; Li, C.; Hu, X.; Zhang, P.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; Choi, Y.; and Gao, J. 2020 · 2004
Earlier work this paper cites.
Improving Vision-and-Language Navigation with Image-Text Pairs from the Web
Majumdar, A.; Shrivastava, A.; Lee, S.; Anderson, P.; Parikh, D.; and Batra, D. 2020 · 2004
Earlier work this paper cites.
lamBERT: Language and Action Learning Using Multimodal BERT
Miyazawa, K.; Aoki, T.; Horii, T.; and Nagai, T. 2020 · 2004
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Matterport3D: Learning from RGB-D Data in Indoor Environments
Chang, A.; Dai, A.; Funkhouser, T.; Halber, M.; Niessner, M.; Savva, M.; Song, S.; Zeng, A.; and Zhang, Y. 2017 · 2017
Cited alongside, same era.
AI2-THOR: An Interactive 3D Environment for Visual AI
Kolve, E.; Mottaghi, R.; Han, W.; VanderBilt, E.; Weihs, L.; Herrasti, A.; Gordon, D.; Zhu, Y.; Gupta, A.; and Farhadi, A. 2017 · 2017
Cited alongside, same era.
Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments
Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sünderhauf, N.; Reid, I.; Gould, S.; and van den Hengel, A. 2018 · 2018
Cited alongside, same era.
Embodied Question Answering
Das, A.; Datta, S.; Gkioxari, G.; Lee, S.; Parikh, D.; and Batra, D. 2018 · 2018
Cited alongside, same era.
Speaker-Follower Models for Vision-and-Language Navigation
Fried, D.; Hu, R.; Cirik, V.; Rohrbach, A.; Andreas, J.; Morency, L.-P.; Berg-Kirkpatrick, T.; Saenko, K.; Klein, D.; and Darrell, T. 2018 · 2018
Cited alongside, same era.
Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation
Jain, V.; Magalhaes, G.; Ku, A.; Vaswani, A.; Ie, E.; and Baldridge, J. 2019 · 2019
Later among the works it cites.
Habitat: A Platform for Embodied AI Research
Savva, M.; Kadian, A.; Maksymets, O.; Zhao, Y.; Wijmans, E.; Jain, B.; Straub, J.; Liu, J.; Koltun, V.; Malik, J.; Parikh, D.; and Batra, D. 2019 · 2019
Later among the works it cites.
Vision-and-Dialog Navigation
Thomason, J.; Murray, M.; Cakmak, M.; and Zettlemoyer, L. 2019 · 2019
Later among the works it cites.
Gibson Env V2: Embodied Simulation Environments for Interactive Navigation
Xia, F.; Li, C.; Chen, K.; Shen, W. B.; Martín-Martín, R.; Hirose, N.; Zamir, A. R.; Fei-Fei, L.; and Savarese, S. 2019 · 2019
Later among the works it cites.
REVERIE: Remote Embodied Visual Referring Expression in Real Indoor Environments
Qi, Y.; Wu, Q.; Anderson, P.; Wang, X.; Wang, W. Y.; Shen, C.; and van den Hengel, A. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chalet: Cornell house agent learning environment
Yan, C.; Misra, D.; Bennnett, A.; Walsman, A.; Bisk, Y.; and Artzi, Y. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Are You Looking? Grounding to Multiple Modalities in Vision-and-Language Navigation
Hu, R.; Fried, D.; Rohrbach, A.; Klein, D.; Darrell, T.; and Saenko, K. 2019 · 2019
Cited alongside, same era.
Shridhar, M.; Thomason, J.; Gordon, D.; Bisk, Y.; Han, W.; Mottaghi, R.; Zettlemoyer, L.; and Fox, D. 2020 · 2020
Later among the works it cites.
Interactive Gibson Benchmark: A Benchmark for Interactive Navigation in Cluttered Environments
Xia, F.; Shen, W. B.; Li, C.; Kasimbeg, P.; Tchapmi, M. E.; Toshev, A.; Martín-Martín, R.; and Savarese, S. 2020 · 2020
Later among the works it cites.
Vision-Language Navigation With Self-Supervised Auxiliary Reasoning Tasks
Zhu, F.; Zhu, Y.; Chang, X.; and Liang, X. 2020 · 2020
Later among the works it cites.