Fetching the paper…
Reading the bibliography…
Building a conversational embodied agent to execute real-life tasks has been a long-standing yet quite challenging research goal, as it requires effective human-agent communication, multi-modal understanding, long-range sequential decision making, etc.
Fast marching methods
Sethian, J. A. 1999 · 1999
Earlier work this paper cites.
Traversability classification using unsupervised on-line visual learning for outdoor robot navigation
Kim, D.; Sun, J.; Oh, S. M.; Rehg, J.; and Bobick, A. 2006 · 2006
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Mask R-CNN
He, K.; Gkioxari, G.; Dollar, P.; and Girshick, R. 2017 · 2017
Earlier work this paper cites.
AI2-THOR: An Interactive 3D Environment for Visual AI
Kolve, E.; Mottaghi, R.; Gordon, D.; Zhu, Y.; Gupta, A.; and Farhadi, A. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Zhu, Y.; Mottaghi, R.; Kolve, E.; Lim, J. J.; Gupta, A.; Fei-Fei, L.; and Farhadi, A. 2017 · 2017
Earlier work this paper cites.
Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments
Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sünderhauf, N.; Reid, I.; Gould, S.; and van den Hengel, A. 2018 · 2018
Earlier work this paper cites.
Speaker-Follower Models for Vision-and-Language Navigation
Fried, D.; Hu, R.; Cirik, V.; Rohrbach, A.; Andreas, J.; Morency, L.-P.; Berg-Kirkpatrick, T.; Saenko, K.; Klein, D.; and Darrell, T. 2018 · 2018
Earlier work this paper cites.
IQA: Visual Question Answering in Interactive Environments
Gordon, D.; Kembhavi, A.; Rastegari, M.; Redmon, J.; Fox, D.; and Farhadi, A. 2018 · 2018
Earlier work this paper cites.
Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction
Misra, D.; Bennett, A.; Blukis, V.; Niklasson, E.; Shatkhin, M.; and Artzi, Y. 2018 · 2018
Earlier work this paper cites.
TOUCHDOWN: Natural Language Navigation and Spatial Reasoning in Visual Street Environments
Chen, H.; Suhr, A.; Misra, D.; Snavely, N.; and Artzi, Y. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Learning by Abstraction: The Neural State Machine
Hudson, D.; and Manning, C. D. 2019 · 2019
Earlier work this paper cites.
The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
Mao, J.; Gan, C.; Kohli, P.; Tenenbaum, J. B.; and Wu, J. 2019 · 2019
Earlier work this paper cites.
Collaborative Dialogue in Minecraft
Narayan-Chen, A.; Jayannavar, P.; and Hockenmaier, J. 2019 · 2019
Cited alongside, same era.
Help, Anna! Visual Navigation with Natural Multimodal Assistance via Retrospective Curiosity-Encouraging Imitation Learning
Nguyen, K.; and Daumé III, H. 2019 · 2019
Cited alongside, same era.
Vision-Based Navigation With Language-Based Assistance via Imitation Learning With Indirect Intervention
Nguyen, K.; Dey, D.; Brockett, C.; and Dolan, B. 2019 · 2019
Cited alongside, same era.
Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout
Tan, H.; Yu, L.; and Bansal, M. 2019 · 2019
Cited alongside, same era.
Vision-and-Dialog Navigation
Thomason, J.; Murray, M.; Cakmak, M.; and Zettlemoyer, L. 2019 · 2019
Cited alongside, same era.
Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
A Persistent Spatial Semantic Representation for High-level Natural Language Instruction Execution
Blukis, V.; Paxton, C.; Fox, D.; Garg, A.; and Artzi, Y. 2021 · 2021
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Later among the works it cites.
Airbert: In-Domain Pretraining for Vision-and-Language Navigation
Guhur, P.-L.; Tapaswi, M.; Chen, S.; Laptev, I.; and Schmid, C. 2021 · 2021
Later among the works it cites.
Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision
He, K.; Huang, Y.; Wu, Q.; Yang, J.; An, D.; Sima, S.; and Wang, L. 2021 · 2021
Later among the works it cites.
VLN BERT: A Recurrent Vision-and-Language BERT for Navigation
Hong, Y.; Wu, Q.; Qi, Y.; Rodriguez-Opazo, C.; and Gould, S. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, X.; Huang, Q.; Celikyilmaz, A.; Gao, J.; Shen, D.; Wang, Y.-F.; Wang, W. Y.; and Zhang, L. 2019 · 2019
Cited alongside, same era.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Scao, T. L.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. M. 2019 · 2019
Cited alongside, same era.
Visual Semantic Navigation using Scene Priors
Yang, W.; Wang, X.; Farhadi, A.; Gupta, A.; and Mottaghi, R. 2019 · 2019
Cited alongside, same era.
The RobotSlang Benchmark: Dialog-guided Robot Localization and Navigation
Banerjee, S.; Thomason, J.; and Corso, J. J. 2020 · 2020
Cited alongside, same era.
Plug and Play Language Models: A Simple Approach to Controlled Text Generation
Dathathri, S.; Madotto, A.; Lan, J.; Hung, J.; Frank, E.; Molino, P.; Yosinski, J.; and Liu, R. 2020 · 2020
Cited alongside, same era.
Neural Module Networks for Reasoning over Text
Gupta, N.; Lin, K.; Roth, D.; Singh, S.; and Gardner, M. 2020 · 2020
Cited alongside, same era.
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding
Ku, A.; Anderson, P.; Patel, R.; Ie, E.; and Baldridge, J. 2020 · 2020
Cited alongside, same era.
NDH-Full: Learning and Evaluating Navigational Agents on Full-Length Dialogue
Kim, H.; Li, J.; and Bansal, M. 2021 · 2021
Later among the works it cites.
TEACh: Task-driven Embodied Agents that Chat
Padmakumar, A.; Thomason, J.; Shrivastava, A.; Lange, P.; Narayan-Chen, A.; Gella, S.; Piramuthu, R.; Tur, G.; and Hakkani-Tur, D. 2021 · 2021
Later among the works it cites.
Episodic Transformer for Vision-and-Language Navigation
Pashevich, A.; Schmid, C.; and Sun, C. 2021 · 2021
Later among the works it cites.
proscript: Partially ordered scripts generation via pre-trained language models
Sakaguchi, K.; Bhagavatula, C.; Bras, R. L.; Tandon, N.; Clark, P.; and Choi, Y. 2021 · 2021
Later among the works it cites.
ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Shridhar, M.; Yuan, X.; Côté, M.-A.; Bisk, Y.; Trischler, A.; and Hausknecht, M. 2021 · 2021
Later among the works it cites.
Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory
Vasudevan, A. B.; Dai, D.; and Van Gool, L. 2021 · 2021
Later among the works it cites.
Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions
Gu, J.; Stefani, E.; Wu, Q.; Thomason, J.; and Wang, X. 2022 · 2022
Closest in time.
Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
Huang, W.; Abbeel, P.; Pathak, D.; and Mordatch, I. 2022 · 2022
Closest in time.
FILM: Following Instructions in Language with Modular Methods
Min, S. Y.; Chaplot, D. S.; Ravikumar, P. K.; Bisk, Y.; and Salakhutdinov, R. 2022 · 2022
Closest in time.
Habitat-Web: Learning Embodied Object-Search Strategies from Human Demonstrations at Scale
Ramrakhya, R.; Undersander, E.; Batra, D.; and Das, A. 2022 · 2022
Closest in time.
One Step at a Time: Long-Horizon Vision-and-Language Navigation with Milestones
Song, C. H.; Kil, J.; Pan, T.-Y.; Sadler, B. M.; Chao, W.-L.; and Su, Y. 2022 · 2022
Closest in time.