Fetching the paper…
Reading the bibliography…
Vision-and-Language Navigation (VLN), as a crucial research problem of Embodied AI, requires an embodied agent to navigate through complex 3D environments following natural language instructions.
P. N. Johnson-Laird, Mental models: Towards a cognitive science of language, inference, and consciousness . Harvard University Press, 1983, no. 6
1983
Earlier work this paper cites.
Johnson-Laird, Philip N, “Mental models and human reasoning,” Proceedings of the National Academy of Sciences , vol. 107, no. 43, pp. 18 243–18 250, 2010
2010
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sunderhauf, I. Reid, S. Gould, and A. van den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in CVPR , 2018
2018
Earlier work this paper cites.
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell, “Speaker-follower models for vision-and-language navigation,” in NeurIPS , 2018
2018
Earlier work this paper cites.
H. Chen, A. Suhr, D. K. Misra, N. Snavely, and Y. Artzi, “Touchdown: Natural language navigation and spatial reasoning in visual street environments,” in CVPR , 2019
2019
Earlier work this paper cites.
V. Jain, G. Magalhaes, A. Ku, A. Vaswani, E. Ie, and J. Baldridge, “Stay on the path: Instruction fidelity in vision-and-language navigation,” in ACL , 2019
2019
Earlier work this paper cites.
H. Tan, L. Yu, and M. Bansal, “Learning to navigate unseen environments: Back translation with environmental dropout,” in NAACL-HLT , 2019
2019
Earlier work this paper cites.
X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y.-F. Wang, W. Y. Wang, and L. Zhang, “Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation,” in CVPR , 2019
2019
Earlier work this paper cites.
C.-Y. Ma, jiasen lu, Z. Wu, G. AlRegib, Z. Kira, richard socher, and C. Xiong, “Self-monitoring navigation agent via auxiliary progress estimation,” in ICLR , 2019
2019
Earlier work this paper cites.
G. Ilharco, V. Jain, A. Ku, E. Ie, and J. Baldridge, “General evaluation for instruction conditioned navigation using dynamic time warping,” ViGIL@NeurIPS , 2019
2019
Earlier work this paper cites.
OpenAI, “Introducing chatgpt,” https://openai.com/blog/chatgpt
2019
Earlier work this paper cites.
Y. Qi, Q. Wu, P. Anderson, X. Wang, W. Y. Wang, C. Shen, and A. van den Hengel, “Reverie: Remote embodied visual referring expression in real indoor environments,” in CVPR , 2020
2020
Earlier work this paper cites.
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge, “Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding,” in EMNLP , 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, and et al., “Language models are few-shot learners,” in NeurIPS , 2020
2020
Earlier work this paper cites.
T.-J. Fu, X. E. Wang, M. F. Peterson, S. T. Grafton, M. P. Eckstein, and W. Y. Wang, “Counterfactual vision-and-language navigation via adversarial path sampling,” in ECCV , 2020
2020
Earlier work this paper cites.
F. Zhu, Y. Zhu, X. Chang, and X. Liang, “Vision-language navigation with self-supervised auxiliary reasoning tasks,” in CVPR , 2020
2020
Earlier work this paper cites.
Z. Deng, K. Narasimhan, and O. Russakovsky, “Evolving graphical planner: Contextual global planning for vision-and-language navigation,” in NeurIPS , 2020
2020
Earlier work this paper cites.
Y. Qi, Z. Pan, S. Zhang, A. van den Hengel, and Q. Wu, “Object-and-action aware model for visual language navigation,” in ECCV , 2020
2020
Earlier work this paper cites.
W. Hao, C. Li, X. Li, L. Carin, and J. Gao, “Towards learning a generic agent for vision-and-language navigation via pre-training,” in CVPR , 2020
2020
Earlier work this paper cites.
C. Liu, F. Zhu, X. Chang, X. Liang, Z. Ge, and Y.-D. Shen, “Vision-language navigation with random environmental mixup,” in ICCV , 2021
2021
Earlier work this paper cites.
B. Lin, Y. Zhu, Y. Long, X. Liang, Q. Ye, and L. Lin, “Adversarial reinforced instruction attacker for robust vision-language navigation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 10, pp. 7175–7189, 2021
2021
Earlier work this paper cites.
Y. Hong, Q. Wu, Y. Qi, C. Rodriguez-Opazo, and S. Gould, “Vln bert: A recurrent vision-and-language bert for navigation,” in CVPR , 2021
2021
Cited alongside, same era.
S. Chen, P.-L. Guhur, C. Schmid, and I. Laptev, “History aware multimodal transformer for vision-and-language navigation,” in NeurIPS , 2021
2021
Cited alongside, same era.
P.-L. Guhur, M. Tapaswi, S. Chen, I. Laptev, and C. Schmid, “Airbert: In-domain pretraining for vision-and-language navigation,” in ICCV , 2021
2021
Cited alongside, same era.
J. Yu, X. Jiang, Z. Qin, W. Zhang, Y. Hu, and Q. Wu, “Learning dual encoding model for adaptive visual understanding in visual dialogue,” IEEE TIP , 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in ICML , 2021
O. OpenAI, “Gpt-4 technical report,” Mar 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
D. An, Y. Qi, Y. Li, Y. Huang, L. Wang, T. Tan, and J. Shao, “Bevbert: Topo-metric map pre-training for language-guided navigation,” in ICCV , 2023
2023
Later among the works it cites.
Z. Wang, J. Li, Y. Hong, Y. Wang, Q. Wu, M. Bansal, S. Gould, H. Tan, and Y. Qiao, “Scaling data generation in vision-and-language navigation,” in ICCV , 2023
2023
Later among the works it cites.
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “React: Synergizing reasoning and acting in language models,” in ICLR , 2023
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in ICML , 2022
2022
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” in NeurIPS , 2022
2022
Cited alongside, same era.
X. Liang, F. Zhu, Y. Zhu, B. Lin, B. Wang, and X. Liang, “Contrastive instruction-trajectory learning for vision-language navigation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 2, 2022, pp. 1592–1600
2022
Cited alongside, same era.
B. Lin, Y. Zhu, Z. Chen, X. Liang, J. Liu, and X. Liang, “Adapt: Vision-language navigation with modality-aligned action prompts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 396–15 406
2022
Cited alongside, same era.
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev, “Think global, act local: Dual-scale graph transformer for vision-and-language navigation,” in CVPR , 2022
2022
Cited alongside, same era.
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid et al. , “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” in CoRL , 2023
2023
Later among the works it cites.
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” in ICLR , 2023
2023
Later among the works it cites.
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. Le et al. , “Least-to-most prompting enables complex reasoning in large language models,” in ICLR , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Long, “Large language model guided tree-of-thought,” arXiv preprint arXiv:2305.08291 , 2023
2023
Later among the works it cites.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Stanford alpaca: An instruction-following llama model,” https://github.com/tatsu-lab/stanford\_alpaca
2023
Later among the works it cites.
2023
Later among the works it cites.
G. Zhou, Y. Hong, and Q. Wu, “Navgpt: Explicit reasoning in vision-and-language navigation with large language models,” in AAAI , 2024
2024
Closest in time.
J. Chen, B. Lin, R. Xu, Z. Chai, X. Liang, and K.-Y. Wong, “Mapgpt: Map-guided prompting with adaptive path planning for vision-and-language navigation,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 9796–9810
2024
Closest in time.
2024
Closest in time.
D. Zheng, S. Huang, L. Zhao, Y. Zhong, and L. Wang, “Towards learning a generalist model for embodied navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 624–13 634
2024
Closest in time.