Fetching the paper…
Reading the bibliography…
There is a growing interest in the community in making an embodied AI agent perform a complicated task while interacting with an environment following natural language directives.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
M. Kempka, M. Wydmuch, G. Runc, J. Toczek, and W. Jaśkowski · 2016
Earlier work this paper cites.
Matterport3d: Learning from rgb-d data in indoor environments
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
AI2-THOR: An Interactive 3D Environment for Visual AI
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi · 2017
Earlier work this paper cites.
Visual semantic planning using deep successor representations
Y. Zhu, D. Gordon, E. Kolve, D. Fox, L. Fei-Fei, A. Gupta, R. Mottaghi, and A. Farhadi · 2017
Earlier work this paper cites.
On evaluation of embodied navigation agents
P. Anderson, A. X. Chang, D. S. Chaplot, A. Dosovitskiy, S. Gupta, V. Koltun, J. Kosecka, J. Malik, R. Mottaghi, M. Savva, and A. R. Zamir · 2018
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. van den Hengel · 2018
Earlier work this paper cites.
Textworld: A learning environment for text-based games
Marc-Alexandre Côté, Ákos Kádár, Xingdi Yuan, Ben Kybartas, Tavian Barnes, Emery Fine, James Moore, Ruo Yu Tao, Matthew Hausknecht, Layla El Asri, Mahmoud Adada, Wendy Tay, and Adam Trischler · 2018
Earlier work this paper cites.
Embodied Question Answering
A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra · 2018
Earlier work this paper cites.
Deep object-centric representations for generalizable robot learning
C. Devin, P. Abbeel, T. Darrell, and S. Levine · 2018
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell · 2018
Cited alongside, same era.
Iqa: Visual question answering in interactive environments
D. Gordon, A. Kembhavi, M. Rastegari, J. Redmon, D. Fox, and A. Farhadi · 2018
Cited alongside, same era.
Virtualhome: Simulating household activities via programs
X. Puig, K. Ra, M. Boben, J. Li, T. Wang, S. Fidler, and A. Torralba · 2018
Cited alongside, same era.
Building generalizable agents with a realistic and rich 3d environment
Yi Wu, Yuxin Wu, Georgia Gkioxari, and Yuandong Tian · 2018
Cited alongside, same era.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
H. Chen, A. Suhr, D. Misra, N. Snavely, and Y. Artzi · 2019
Cited alongside, same era.
What should i do now? marrying reinforcement learning and symbolic planning
Modularity improves out-of-domain instruction following
R. Corona, D. Fried, C. Devin, D. Klein, and T. Darrell · 2020
Later among the works it cites.
Learning to follow directions in street view
K. M. Hermann, M. Malinowski, P. Mirowski, A. Banki-Horvath, K. Anderson, and R. Hadsell · 2020
Later among the works it cites.
Beyond the nav-graph: Vision-and-language navigation in continuous environments
J. Krantz, E. Wijmans, A. Majumdar, D. Batra, and S. Lee · 2020
Later among the works it cites.
Improving vision-and-language navigation with image-text pairs from the web
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra · 2020
Later among the works it cites.
Efficient attention mechanism for visual dialog that can handle all the interactions between multiple inputs
V. Q. Nguyen, M. Suganuma, and T. Okatani · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Gordon, D. Fox, and A. Farhadi · 2019
Cited alongside, same era.
Self-monitoring navigation agent via auxiliary progress estimation
C.-Y. Ma, J. Lu, Z. Wu, G. AlRegib, Z. Kira, R. Socher, and C. Xiong · 2019
Cited alongside, same era.
Vision-based navigation with language-based assistance via imitation learning with indirect intervention
K. Nguyen, D. Dey, C. Brockett, and B. Dolan · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Cited alongside, same era.
Executing instructions in situated collaborative interactions
Alane Suhr, Claudia Yan, Jack Schluger, Stanley Yu, Hadi Khader, Marwa Mouallem, Iris Zhang, and Yoav Artzi · 2019
Cited alongside, same era.
Learning to navigate unseen environments: Back translation with environmental dropout
H. Tan, L. Yu, and M. Bansal · 2019
Cited alongside, same era.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y.-F. Wang, W. Y. Wang, and L. Zhang · 2019
Cited alongside, same era.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox · 2020
Later among the works it cites.
Improving mask prediction for long horizon instruction following
K. P. Singh, S. Bhambri, B. Kim, , and J. Choi · 2020
Later among the works it cites.
Moca: A modular object-centric approach for interactive instruction following
K. P. Singh, S. Bhambri, B. Kim, R. Mottaghi, and J. Choi · 2020
Later among the works it cites.
Vision-and-dialog navigation
J. Thomason, M. Murray, M. Cakmak, and L. Zettlemoyer · 2020
Later among the works it cites.
Alfred speaks: Automatic instruction generation for egocentric skill learning
L. Yeung, Y. Bisk, and O. Polozov · 2020
Later among the works it cites.
Vision-language navigation with self-supervised auxiliary reasoning tasks
F. Zhu, Y. Zhu, X. Chang, and X. Liang · 2020
Later among the works it cites.
{ALFW}orld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Cote, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht · 2021
Closest in time.