Fetching the paper…
Reading the bibliography…
Recent research on Vision-and-Language Navigation (VLN) indicates that agents suffer from poor generalization in unseen environments due to the lack of realistic training environments and high-quality path-instruction pairs.
A. X. Chang, A. Dai, T. A. Funkhouser, M. Halber, M. Nießner, M. Savva, S. Song, A. Zeng, and Y. Zhang, “Matterport3d: Learning from RGB-D data in indoor environments,” in 3DV , 2017, pp. 667–676
2017
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 3674–3683
2018
Earlier work this paper cites.
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell, “Speaker-follower models for vision-and-language navigation,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
H. Tan, L. Yu, and M. Bansal, “Learning to navigate unseen environments: Back translation with environmental dropout,” in NAACL , J. Burstein, C. Doran, and T. Solorio, Eds., 2019, pp. 2610–2621
2019
Earlier work this paper cites.
C.-Y. Ma, Z. Wu, G. AlRegib, C. Xiong, and Z. Kira, “The regretful agent: Heuristic-aided navigation through progress estimation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6732–6740
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Ma, J. Lu, Z. Wu, G. AlRegib, Z. Kira, R. Socher, and C. Xiong, “Self-monitoring navigation agent via auxiliary progress estimation,” in ICLR , 2019
2019
Earlier work this paper cites.
H. Chen, A. Suhr, D. Misra, N. Snavely, and Y. Artzi, “Touchdown: Natural language navigation and spatial reasoning in visual street environments,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 12 538–12 547
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y.-F. Wang, W. Y. Wang, and L. Zhang, “Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6629–6638
2019
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
Y. Zhang, H. Tan, and M. Bansal, “Diagnosing the environment bias in vision-and-language navigation,” 2020. [Online]. Available: https://openreview.net/forum?id=S1eYKlrYvr
2020
Earlier work this paper cites.
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra, “Improving vision-and-language navigation with image-text pairs from the web,” in ECCV , 2020, pp. 259–274
2020
Cited alongside, same era.
M. Chang, A. Gupta, and S. Gupta, “Semantic visual navigation by watching youtube videos,” Advances in Neural Information Processing Systems , vol. 33, pp. 4283–4294, 2020
2020
Cited alongside, same era.
J. Thomason, M. Murray, M. Cakmak, and L. Zettlemoyer, “Vision-and-dialog navigation,” in Conference on Robot Learning . PMLR, 2020, pp. 394–406
2020
Cited alongside, same era.
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge, “Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding,” in EMNLP , 2020, pp. 4392–4412
2020
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Later among the works it cites.
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev, “Learning from unlabeled 3d environments for vision-and-language navigation,” in European Conference on Computer Vision . Springer, 2022, pp. 638–655
2022
Later among the works it cites.
J. Gu, E. Stefani, Q. Wu, J. Thomason, and X. Wang, “Vision-and-language navigation: A survey of tasks, methods, and future directions,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, 2022
2022
Later among the works it cites.
J. Li, H. Tan, and M. Bansal, “Envedit: Environment editing for vision-and-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 407–15 417
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Qi, Q. Wu, P. Anderson, X. Wang, W. Y. Wang, C. Shen, and A. v. d. Hengel, “Reverie: Remote embodied visual referring expression in real indoor environments,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 9982–9991
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Parvaneh, E. Abbasnejad, D. Teney, J. Q. Shi, and A. Van den Hengel, “Counterfactual vision-and-language navigation: Unravelling the unseen,” Advances in neural information processing systems , vol. 33, pp. 5296–5307, 2020
2020
Cited alongside, same era.
P.-L. Guhur, M. Tapaswi, S. Chen, I. Laptev, and C. Schmid, “Airbert: In-domain pretraining for vision-and-language navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1634–1643
2021
Cited alongside, same era.
K. He, Y. Huang, Q. Wu, J. Yang, D. An, S. Sima, and L. Wang, “Landmark-rxr: Solving vision-and-language navigation with fine-grained alignment supervision,” Advances in Neural Information Processing Systems , vol. 34, pp. 652–663, 2021
2021
Cited alongside, same era.
H. Kim, J. Li, and M. Bansal, “Ndh-full: Learning and evaluating navigational agents on full-length dialogue,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021
2021
Cited alongside, same era.
C. Liu, F. Zhu, X. Chang, X. Liang, Z. Ge, and Y.-D. Shen, “Vision-language navigation with random environmental mixup,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1644–1654
2021
Cited alongside, same era.
2022
Later among the works it cites.
X. Liang, F. Zhu, L. Li, H. Xu, and X. Liang, “Visual-language navigation pretraining via prompt-based environmental self-exploration,” in ACL , 2022, pp. 4837–4851
2022
Later among the works it cites.
S. Chen, P. Guhur, M. Tapaswi, C. Schmid, and I. Laptev, “Think global, act local: Dual-scale graph transformer for vision-and-language navigation,” in CVPR , 2022, pp. 16 516–16 526
2022
Later among the works it cites.
A. Kamath, P. Anderson, S. Wang, J. Y. Koh, A. Ku, A. Waters, Y. Yang, J. Baldridge, and Z. Parekh, “A new path: Scaling vision-and-language navigation with synthetic instructions and imitation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 10 813–10 823
2023
Later among the works it cites.
K. Lin, P. Chen, D. Huang, T. H. Li, M. Tan, and C. Gan, “Learning vision-and-language navigation from youtube videos,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 8317–8326
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
A. Padmakumar, J. Thomason, A. Shrivastava, P. Lange, A. Narayan-Chen, S. Gella, R. Piramuthu, G. Tur, and D. Hakkani-Tur, “Teach: Task-driven embodied agents that chat,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 2, 2022, pp. 2017–2025
2025
Closest in time.