Fetching the paper…
Reading the bibliography…
Vision-and-Language Navigation in Continuous Environments (VLN-CE) is one of the most intuitive yet challenging embodied AI tasks.
S. Bird and E. Loper, “NLTK: the natural language toolkit,” in ACL , 2004. http://dx.doi.org/10.3115/1219044.1219075
2004
Earlier work this paper cites.
J. Hort, J. Laczó, M. Vyhnálek, M. Bojar, J. Bureš, and K. Vlček, “Spatial navigation deficit in amnestic mild cognitive impairment,” Proc. of the National Academy of Sciences , vol. 104, no. 10, p. 4042–4047, Mar. 2007. http://dx.doi.org/10.1073/pnas.0611314104
2007
Earlier work this paper cites.
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niebner, M. Savva, S. Song, A. Zeng, and Y. Zhang, “Matterport3D: Learning from RGB-D Data in Indoor Environments,” in 3DV , CA, USA, oct 2017, pp. 667–676. https://doi.ieeecomputersociety.org/10.1109/3DV.2017.00081
2017
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sunderhauf, I. Reid, S. Gould, and A. van den Hengel, “Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments,” in CVPR , Jun. 2018. http://dx.doi.org/10.1109/cvpr.2018.00387
2018
Earlier work this paper cites.
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell, “Speaker-Follower Models for Vision-and-Language Navigation,” in NeurIPS , vol. 31. Curran Associates, Inc., 2018. https://proceedings.neurips.cc/paper_files/paper/2018/file/6a81681a7af700c6385d36577ebec359-Paper.pdf
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, D. Parikh, and D. Batra, “Habitat: A Platform for Embodied AI Research,” in ICCV , Oct. 2019. http://dx.doi.org/10.1109/iccv.2019.00943
2019
Earlier work this paper cites.
X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y.-F. Wang, W. Y. Wang, and L. Zhang, “Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation,” in CVPR , Jun. 2019. http://dx.doi.org/10.1109/cvpr.2019.00679
2019
Earlier work this paper cites.
H. Huang, V. Jain, H. Mehta, A. Ku, G. Magalhaes, J. Baldridge, and E. Ie, “Transferable Representation Learning in Vision-and-Language Navigation,” in ICCV , Oct. 2019. http://dx.doi.org/10.1109/iccv.2019.00750
2019
Earlier work this paper cites.
H. Tan, L. Yu, and M. Bansal, “Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout,” in NAACL , Minneapolis, Minnesota, Jun. 2019, pp. 2610–2621. https://aclanthology.org/N19-1268
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , vol. 1, 2019, p. 4171 – 4186. https://www.scopus.com/inward/record.uri?eid=2-s2.0-85083815650&partnerID=40&md5=4986c6d6076c0c91df84d17216b47216
2019
Earlier work this paper cites.
H. Tan and M. Bansal, “”LXMERT: Learning cross-modality encoder representations from transformers”,” in EMNLP-IJCNLP , K. Inui, J. Jiang, V. Ng, and X. Wan, Eds., Hong Kong, China, Nov. 2019, pp. 5100–5111. https://aclanthology.org/D19-1514
2019
Cited alongside, same era.
J. Krantz, E. Wijmans, A. Majumdar, D. Batra, and S. Lee, Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments . ECCV, 2020, p. 104–120. http://dx.doi.org/10.1007/978-3-030-58604-1_7
2020
Cited alongside, same era.
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge, “Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding,” in EMNLP , 2020. http://dx.doi.org/10.18653/v1/2020.emnlp-main.356
2020
Cited alongside, same era.
Y. Zhang, H. Tan, and M. Bansal, “Diagnosing the Environment Bias in Vision-and-Language Navigation,” in IJCAI , ser. IJCAI-PRICAI-2020, Jul. 2020. http://dx.doi.org/10.24963/ijcai.2020/124
2020
Cited alongside, same era.
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev, “Think Global, Act Local: Dual-Scale Graph Transformer for Vision-and-Language Navigation,” in CVPR , Jun. 2022. http://dx.doi.org/10.1109/cvpr52688.2022.01604
2022
Later among the works it cites.
W. Zhu, Y. Qi, P. Narayana, K. Sone, S. Basu, X. Wang, Q. Wu, M. Eckstein, and W. Y. Wang, “Diagnosing Vision-and-Language Navigation: What Really Matters,” in NAACL , Seattle, United States, Jul. 2022, pp. 5981–5993. https://aclanthology.org/2022.naacl-main.438
2022
Later among the works it cites.
X. Liang, F. Zhu, Y. Zhu, B. Lin, B. Wang, and X. Liang, “Contrastive Instruction-Trajectory Learning for Vision-Language Navigation,” AAAI , vol. 36, no. 2, pp. 1592–1600, Jun. 2022. https://ojs.aaai.org/index.php/AAAI/article/view/20050
2022
Later among the works it cites.
Z. Wang, J. Li, Y. Hong, Y. Wang, Q. Wu, M. Bansal, S. Gould, H. Tan, and Y. Qiao, “Scaling Data Generation in Vision-and-Language Navigation,” in ICCV , Oct. 2023. http://dx.doi.org/10.1109/iccv51070.2023.01103
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Chen, P.-L. Guhur, C. Schmid, and I. Laptev, “History Aware Multimodal Transformer for Vision-and-Language Navigation,” in NeurIPS , vol. 34. Curran Associates, Inc., 2021, pp. 5834–5847. https://proceedings.neurips.cc/paper_files/paper/2021/file/2e5c2cb8d13e8fba78d95211440ba326-Paper.pdf
2021
Cited alongside, same era.
Y. Hong, Q. Wu, Y. Qi, C. Rodriguez-Opazo, and S. Gould, “VLN ↺ \circlearrowleft BERT: A Recurrent Vision-and-Language BERT for Navigation,” in CVPR , Jun. 2021. http://dx.doi.org/10.1109/cvpr46437.2021.00169
2021
Cited alongside, same era.
M. Zhao, P. Anderson, V. Jain, S. Wang, A. Ku, J. Baldridge, and E. Ie, “On the Evaluation of Vision-and-Language Navigation Instructions,” in EACL , Online, Apr. 2021, pp. 1302–1316. https://aclanthology.org/2021.eacl-main.111
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning Transferable Visual Models From Natural Language Supervision,” in ICML , vol. 139. PMLR, 18–24 Jul 2021, pp. 8748–8763. https://proceedings.mlr.press/v139/radford21a.html
2021
Cited alongside, same era.
J. Gu, E. Stefani, Q. Wu, J. Thomason, and X. Wang, “Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions,” in Proc. of Association for Computational Linguistics . ACL, 2022. http://dx.doi.org/10.18653/v1/2022.acl-long.524
2022
Cited alongside, same era.
Y. Hong, Z. Wang, Q. Wu, and S. Gould, “Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation,” in CVPR , Jun. 2022. http://dx.doi.org/10.1109/cvpr52688.2022.01500
2022
Cited alongside, same era.
2023
Later among the works it cites.
D. An, Y. Qi, Y. Li, Y. Huang, L. Wang, T. Tan, and J. Shao, “BEVBert: Multimodal Map Pre-training for Language-guided Navigation,” ICCV , 2023. https://doi.org/10.48550/arXiv.2212.04385
2023
Later among the works it cites.
Z. Wang, X. Li, J. Yang, Y. Liu, and S. Jiang, “GridMM: Grid Memory Map for Vision-and-Language Navigation,” in ICCV , oct 2023, pp. 15 579–15 590. https://doi.ieeecomputersociety.org/10.1109/ICCV51070.2023.01432
2023
Later among the works it cites.
M. Hahn, A. Raj, and J. M. Rehg, “Which way is ‘right’?: Uncovering limitations of Vision-and-Language Navigation models,” in AAMAS , Richland, SC, 2023, p. 2415–2417. https://api.semanticscholar.org/CorpusID:253381693
2023
Later among the works it cites.
Z. Yang, A. Majumdar, and S. Lee, “Behavioral Analysis of Vision-and-Language Navigation Agents,” in CVPR , Jun. 2023. http://dx.doi.org/10.1109/cvpr52729.2023.00253
2023
Later among the works it cites.
D. An, H. Wang, W. Wang, Z. Wang, Y. Huang, K. He, and L. Wang, “ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments,” TPAMI , pp. 1–16, 2024
2024
Closest in time.
F. Taioli, S. Rosa, A. Castellini, L. Natale, A. Del Bue, A. Farinelli, M. Cristani, and Y. Wang, “I2EDL: Interactive Instruction Error Detection and Localization,” in 2024 33rd IEEE International Conference on Robot and Human Interactive Communication (ROMAN) , 2024, pp. 1872–1877. https://doi.org/10.1109/RO-MAN60168.2024.10731349
2024
Closest in time.