Fetching the paper…
Reading the bibliography…
Robots operating in human spaces must be able to engage in natural language interaction with people, both understanding and executing instructions, and using conversation to resolve ambiguity and recover from mistakes.
The Symbol Grounding Problem
Harnad, S. 1990 · 1990
Earlier work this paper cites.
PDDL The Planning Domain Definition Language
Ghallab, M.; Howe, A.; Knoblock, C.; McDermott, D.; Ram, A.; Veloso, M.; Weld, D.; and Wilkins, D. 1998 · 1998
Earlier work this paper cites.
Walk the Talk: Connecting Language, Knowledge, and Action in Route Instructions
MacMahon, M.; Stankiewicz, B.; and Kuipers, B. 2006 · 2006
Earlier work this paper cites.
Learning to Interpret Natural Language Navigation Instructions from Observations
Chen, D.; and Mooney, R. 2011 · 2011
Earlier work this paper cites.
Generalized Grounding Graphs: A Probabilistic Framework for Understanding Grounded Language
Kollar, T.; Tellex, S.; Walter, M. R.; Huang, A.; Bachrach, A.; Hemachandra, S.; Brunskill, E.; Banerjee, A.; Roy, D.; Teller, S.; et al. 2013 · 2013
Earlier work this paper cites.
Learning to Parse Natural Language Commands to a Robot Control System
Matuszek, C.; Herbst, E.; Zettlemoyer, L.; and Fox, D. 2013 · 2013
Earlier work this paper cites.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
Mei, H.; Bansal, M.; and Walter, M. 2016 · 2016
Earlier work this paper cites.
Asking for Help Using Inverse Semantics
Tellex, S.; Knepper, R. A.; Li, A.; Roy, N.; and Rus, D. 2016 · 2016
Earlier work this paper cites.
Mask R-CNN
He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. B. 2017 · 2017
Earlier work this paper cites.
AI2-THOR: An Interactive 3D Environment for Visual AI
Kolve, E.; Mottaghi, R.; Han, W.; VanderBilt, E.; Weihs, L.; Herrasti, A.; Gordon, D.; Zhu, Y.; Gupta, A.; and Farhadi, A. 2017 · 2017
Earlier work this paper cites.
Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments
Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sünderhauf, N.; Reid, I.; Gould, S.; and van den Hengel, A. 2018 · 2018
Earlier work this paper cites.
Grounding Natural Language Instructions to Semantic Goal Representations for Abstraction and Generalization
Arumugam, D.; Karamcheti, S.; Gopalan, N.; Williams, E. C.; Rhee, M.; Wong, L. L.; and Tellex, S. 2018 · 2018
Earlier work this paper cites.
Learning Interpretable Spatial Operations in a Rich 3D Blocks World
Bisk, Y.; Shih, K.; Choi, Y.; and Marcu, D. 2018 · 2018
Earlier work this paper cites.
Mapping Navigation Instructions to Continuous Control Actions with Position Visitation Prediction
Blukis, V.; Misra, D.; Knepper, R. A.; and Artzi, Y. 2018 · 2018
Cited alongside, same era.
Speaker-Follower Models for Vision-and-Language Navigation
Fried, D.; Hu, R.; Cirik, V.; Rohrbach, A.; Andreas, J.; Morency, L.-P.; Berg-Kirkpatrick, T.; Saenko, K.; Klein, D.; and Darrell, T. 2018 · 2018
Cited alongside, same era.
Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction
Misra, D. K.; Bennett, A.; Blukis, V.; Niklasson, E.; Shatkhin, M.; and Artzi, Y. 2018 · 2018
Cited alongside, same era.
VirtualHome: Simulating Household Activities via Programs
Puig, X.; Ra, K.; Boben, M.; Li, J.; Wang, T.; Fidler, S.; and Torralba, A. 2018 · 2018
Cited alongside, same era.
Shifting the Baseline: Single Modality Performance on Visual Navigation & QA
Thomason, J.; Gordon, D.; and Bisk, Y. 2018 · 2018
Cited alongside, same era.
ArraMon: A Joint Navigation-Assembly Instruction Interpretation Task in Dynamic Environments
Kim, H.; Zala, A.; Burri, G.; Tan, H.; and Bansal, M. 2020 · 2020
Later among the works it cites.
RMM: A Recursive Mental Model for Dialog Navigation
Roman, H. R.; Bisk, Y.; Thomason, J.; Celikyilmaz, A.; and Gao, J. 2020 · 2020
Later among the works it cites.
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Shridhar, M.; Thomason, J.; Gordon, D.; Bisk, Y.; Han, W.; Mottaghi, R.; Zettlemoyer, L.; and Fox, D. 2020 · 2020
Later among the works it cites.
Jointly improving parsing and perception for natural language commands through human-robot dialog
Thomason, J.; Padmakumar, A.; Sinapov, J.; Walker, N.; Jiang, Y.; Yedidsion, H.; Hart, J.; Stone, P.; and Mooney, R. 2020 · 2020
Later among the works it cites.
BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby Steps
Zhu, W.; Hu, H.; Chen, J.; Deng, Z.; Jain, V.; Ie, E.; and Sha, F. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Touchdown: Natural Language Navigation and Spatial Reasoning in Visual Street Environments
Chen, H.; Suhr, A.; Misra, D.; Snavely, N.; and Artzi, Y. 2019 · 2019
Cited alongside, same era.
Collaborative Dialogue in Minecraft
Narayan-Chen, A.; Jayannavar, P.; and Hockenmaier, J. 2019 · 2019
Cited alongside, same era.
Help, Anna! Visual Navigation with Natural Multimodal Assistance via Retrospective Curiosity-Encouraging Imitation Learning
Nguyen, K.; and Daumé III, H. 2019 · 2019
Cited alongside, same era.
Executing Instructions in Situated Collaborative Interactions
Suhr, A.; Yan, C.; Schluger, J.; Yu, S.; Khader, H.; Mouallem, M.; Zhang, I.; and Artzi, Y. 2019 · 2019
Cited alongside, same era.
Vision-and-Dialog Navigation
Thomason, J.; Murray, M.; Cakmak, M.; and Zettlemoyer, L. 2019 · 2019
Cited alongside, same era.
Imitating Interactive Intelligence
Abramson, J.; Ahuja, A.; Brussee, A.; Carnevale, F.; Cassin, M.; Clark, S.; Dudzik, A.; Georgiev, P.; Guy, A.; Harley, T.; Hill, F.; Hung, A.; Kenton, Z.; Landon, J.; Lillicrap, T.; Mathewson, K.; Muldal, A.; Santoro, A.; Savinov, N.; Varma, V.; Wayne, G.; Wong, N.; Yan, C.; and Zhu, R. 2020 · 2020
Cited alongside, same era.
Experience Grounds Language
Bisk, Y.; Holtzman, A.; Thomason, J.; Andreas, J.; Bengio, Y.; Chai, J.; Lapata, M.; Lazaridou, A.; May, J.; Nisnevich, A.; Pinto, N.; and Turian, J. 2020 · 2020
Cited alongside, same era.
A Persistent Spatial Semantic Representation for High-level Natural Language Instruction Execution
Blukis, V.; Paxton, C.; Fox, D.; Garg, A.; and Artzi, Y. 2021 · 2021
Closest in time.
Agent with the Big Picture: Perceiving Surroundings for Interactive Instruction Following
Kim, B.; Bhambri, S.; Singh, K. P.; Mottaghi, R.; and Choi, J. 2021 · 2021
Closest in time.
Episodic Transformer for Vision-and-Language Navigation
Pashevich, A.; Schmid, C.; and Sun, C. 2021 · 2021
Closest in time.
The MineRL BASALT Competition on Learning from Human Feedback
Shah, R.; Wild, C.; Wang, S. H.; Alex, N.; Houghton, B.; Guss, W.; Mohanty, S.; Kanervisto, A.; Milani, S.; Topin, N.; et al. 2021 · 2021
Closest in time.
VISITRON: Visual Semantics-Aligned Interactively Trained Object-Navigator
Shrivastava, A.; Gopalakrishnan, K.; Liu, Y.; Piramuthu, R.; Tür, G.; Parikh, D.; and Hakkani-Tür, D. 2021 · 2021
Closest in time.
Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
Suglia, A.; Gao, Q.; Thomason, J.; Thattai, G.; and Sukhatme, G. 2021 · 2021
Closest in time.
Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring
Zhang, Y.; and Chai, J. 2021 · 2021
Closest in time.
LUMINOUS: Indoor Scene Generation for Embodied AI Challenges
Zhao, Y.; Lin, K.; Jia, Z.; Gao, Q.; Thattai, G.; Thomason, J.; and Sukhatme, G. S. 2021 · 2021
Closest in time.