Fetching the paper…
Reading the bibliography…
Vision-and-Language Navigation (VLN) tasks require an agent to follow textual instructions to navigate through 3D environments.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Annual Meeting of the Association for Computational Linguistics , 2002
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Annual Meeting of the Association for Computational Linguistics , 2004
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in IEEvaluation@ACL , 2005
2005
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conf. Comput. Vis. Pattern Recog. IEEE Computer Society, 2016, pp. 770–778
2016
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “SPICE: semantic propositional image caption evaluation,” in Eur. Conf. Comput. Vis. , 2016, pp. 382–398
2016
Earlier work this paper cites.
A. X. Chang, A. Dai, T. A. Funkhouser, M. Halber, M. Nießner, M. Savva, S. Song, A. Zeng, and Y. Zhang, “Matterport3d: Learning from RGB-D data in indoor environments,” in International Conference on 3D Vision , 2017, pp. 667–676
2017
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. D. Reid, S. Gould, and A. van den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2018, pp. 3674–3683
2018
Earlier work this paper cites.
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell, “Speaker-follower models for vision-and-language navigation,” in Adv. Neural Inform. Process. Syst. , 2018, pp. 3318–3329
2018
Earlier work this paper cites.
C. Ma, Z. Wu, G. AlRegib, C. Xiong, and Z. Kira, “The regretful agent: Heuristic-aided navigation through progress estimation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 6732–6740
2019
Earlier work this paper cites.
X. Wang, Q. Huang, A. Çelikyilmaz, J. Gao, D. Shen, Y. Wang, W. Y. Wang, and L. Zhang, “Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 6629–6638
2019
Earlier work this paper cites.
H. Tan, L. Yu, and M. Bansal, “Learning to navigate unseen environments: Back translation with environmental dropout,” in NAACL-HLT , 2019, pp. 2610–2621
2019
Earlier work this paper cites.
H. Huang, V. Jain, H. Mehta, A. Ku, G. Magalhães, J. Baldridge, and E. Ie, “Transferable representation learning in vision-and-language navigation,” in Int. Conf. Comput. Vis. , 2019, pp. 7403–7412
2019
Earlier work this paper cites.
J. Krantz, E. Wijmans, A. Majumdar, D. Batra, and S. Lee, “Beyond the nav-graph: Vision-and-language navigation in continuous environments,” in Eur. Conf. Comput. Vis. , A. Vedaldi, H. Bischof, T. Brox, and J. Frahm, Eds., 2020, pp. 104–120
2020
Earlier work this paper cites.
F. Zhu, Y. Zhu, X. Chang, and X. Liang, “Vision-language navigation with self-supervised auxiliary reasoning tasks,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2020, pp. 10 009–10 019
2020
Earlier work this paper cites.
Y. Qi, Z. Pan, S. Zhang, A. van den Hengel, and Q. Wu, “Object-and-action aware model for visual language navigation,” in Eur. Conf. Comput. Vis. , 2020, pp. 303–317
2020
Earlier work this paper cites.
Y. Hong, C. R. Opazo, Y. Qi, Q. Wu, and S. Gould, “Language and visual entity relationship graph for agent navigation,” in Adv. Neural Inform. Process. Syst. , 2020
2020
Earlier work this paper cites.
W. Hao, C. Li, X. Li, L. Carin, and J. Gao, “Towards learning a generic agent for vision-and-language navigation via pre-training,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2020, pp. 13 134–13 143
2020
Cited alongside, same era.
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra, “Improving vision-and-language navigation with image-text pairs from the web,” in Eur. Conf. Comput. Vis. , 2020, pp. 259–274
2020
Cited alongside, same era.
J. Krantz, A. Gokaslan, D. Batra, S. Lee, and O. Maksymets, “Waypoint models for instruction-guided navigation in continuous environments,” Int. Conf. Comput. Vis. , pp. 15 142–15 151, 2021
2021
Cited alongside, same era.
P. Guhur, M. Tapaswi, S. Chen, I. Laptev, and C. Schmid, “Airbert: In-domain pretraining for vision-and-language navigation,” in Int. Conf. Comput. Vis. , 2021, pp. 1634–1643
2021
Cited alongside, same era.
Y. Long, X. Li, W. Cai, and H. Dong, “Discuss before moving: Visual language navigation via multi-expert discussions,” IEEE International Conference on Robotics and Automation , pp. 17 380–17 387, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Li, Z. Wang, J. Yang, Y. Wang, and S. Jiang, “Kerm: Knowledge enhanced reasoning for vision-and-language navigation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2023, pp. 2583–2592
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Chen, P. Guhur, C. Schmid, and I. Laptev, “History aware multimodal transformer for vision-and-language navigation,” in Adv. Neural Inform. Process. Syst. , 2021, pp. 5834–5847
2021
Cited alongside, same era.
D. An, Y. Qi, Y. Huang, Q. Wu, L. Wang, and T. Tan, “Neighbor-view enhanced model for vision and language navigation,” in Proceedings of the ACM International Conference on Multimedia , 2021
2021
Cited alongside, same era.
J. Li, H. Tan, and M. Bansal, “Envedit: Environment editing for vision-and-language navigation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 15 386–15 396
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems , 2022
2022
Cited alongside, same era.
Y. Hong, Z. Wang, Q. Wu, and S. Gould, “Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 15 418–15 428
2022
Cited alongside, same era.
2022
Cited alongside, same era.
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev, “Think global, act local: Dual-scale graph transformer for vision-and-language navigation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022
2022
Cited alongside, same era.
Y. Qiao, Y. Qi, Y. Hong, Z. Yu, P. Wang, and Q. Wu, “Hop+: History-enhanced and order-aware pre-training for vision-and-language navigation,” IEEE Trans. Pattern Anal. Mach. Intell. , 2023
2023
Later among the works it cites.
D. An, H. Wang, W. Wang, Z. Wang, Y. Huang, K. He, and L. Wang, “Etpnav: Evolving topological planning for vision-language navigation in continuous environments,” IEEE Transactions on pattern analysis and machine intelligence , vol. PP, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:257985276
2023
Later among the works it cites.
D. An, Y. Qi, Y. Li, Y. Huang, L. Wang, T. Tan, and J. Shao, “Bevbert: Multimodal map pre-training for language-guided navigation,” in Int. Conf. Comput. Vis. , 2023, pp. 2737–2748
2023
Later among the works it cites.
Z. Wang, X. Li, J. Yang, Y. Liu, and S. Jiang, “Gridmm: Grid memory map for vision-and-language navigation,” in Int. Conf. Comput. Vis. , 2023, pp. 15 625–15 636
2023
Later among the works it cites.
Y. Qiao, Y. Qi, Z. Yu, J. Liu, and Q. Wu, “March in chat: Interactive prompting for remote embodied referring expression,” in Int. Conf. Comput. Vis. , 2023, pp. 15 758–15 767
2023
Later among the works it cites.
2023
Later among the works it cites.
D. Zheng, S. Huang, L. Zhao, Y. Zhong, and L. Wang, “Towards learning a generalist model for embodied navigation,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2024, pp. 13 624–13 634
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Long, W. Cai, H. Wang, G. Zhan, and H. Dong, “Instructnav: Zero-shot system for generic instruction navigation in unexplored environment,” ArXiv , 2024
2024
Closest in time.