Fetching the paper…
Reading the bibliography…
Vision-and-language navigation (VLN) enables the agent to navigate to a remote location in 3D environments following the natural language instruction.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Matterport3d: Learning from rgb-d data in indoor environments
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niebner, M. Savva, S. Song, A. Zeng, and Y. Zhang · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel · 2018
Earlier work this paper cites.
Habitat: A platform for embodied ai research
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, et al · 2019
Earlier work this paper cites.
Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge · 2020
Earlier work this paper cites.
Reverie: Remote embodied visual referring expression in real indoor environments
Y. Qi, Q. Wu, P. Anderson, X. Wang, W. Y. Wang, C. Shen, and A. v. d. Hengel · 2020
Earlier work this paper cites.
Beyond the nav-graph: Vision-and-language navigation in continuous environments
J. Krantz, E. Wijmans, A. Majumdar, D. Batra, and S. Lee · 2020
Earlier work this paper cites.
Sim-to-real transfer for vision-and-language navigation
P. Anderson, A. Shrivastava, J. Truong, A. Majumdar, D. Parikh, D. Batra, and S. Lee · 2021
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
History aware multimodal transformer for vision-and-language navigation
S. Chen, P.-L. Guhur, C. Schmid, and I. Laptev · 2021
Earlier work this paper cites.
Hierarchical object-to-zone graph for object navigation
S. Zhang, X. Song, Y. Bai, W. Li, Y. Chu, and S. Jiang · 2021
Earlier work this paper cites.
Semantic mapnet: Building allocentric semantic maps and representations from egocentric views
V. Cartillier, Z. Ren, N. Jain, S. Lee, I. Essa, and D. Batra · 2021
Earlier work this paper cites.
Weakly-supervised multi-granularity map learning for vision-and-language navigation
P. Chen, D. Ji, K. Lin, R. Zeng, T. H. Li, M. Tan, and C. Gan · 2022
Cited alongside, same era.
Cross-modal map learning for vision and language navigation
G. Georgakis, K. Schmeckpeper, K. Wanchoo, S. Dan, E. Miltsakaki, D. Roth, and K. Daniilidis · 2022
Cited alongside, same era.
Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation
Y. Hong, Z. Wang, Q. Wu, and S. Gould · 2022
Cited alongside, same era.
Think global, act local: Dual-scale graph transformer for vision-and-language navigation
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev · 2022
Cited alongside, same era.
Poni: Potential functions for objectgoal navigation with interaction-free learning
S. K. Ramakrishnan, D. S. Chaplot, Z. Al-Halah, J. Malik, and K. Grauman · 2022
Cited alongside, same era.
Visual language maps for robot navigation
C. Huang, O. Mees, A. Zeng, and W. Burgard · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, et al · 2023
Later among the works it cites.
Lerf: Language embedded radiance fields
J. Kerr, C. M. Kim, K. Goldberg, A. Kanazawa, and M. Tancik · 2023
Later among the works it cites.
P. Chen, X. Sun, H. Zhi, R. Zeng, T. H. Li, G. Liu, M. Tan, and C. Gan · 2023
Later among the works it cites.
Esceme: Vision-and-language navigation with episodic scene memory
Q. Zheng, D. Liu, C. Wang, J. Zhang, D. Wang, and D. Tao · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generative meta-adversarial network for unseen object navigation
S. Zhang, W. Li, X. Song, Y. Bai, and S. Jiang · 2022
Cited alongside, same era.
Sim-2-sim transfer for vision-and-language navigation in continuous environments
J. Krantz and S. Lee · 2022
Cited alongside, same era.
Bevbert: Multimodal map pre-training for language-guided navigation
D. An, Y. Qi, Y. Li, Y. Huang, L. Wang, T. Tan, and J. Shao · 2023
Cited alongside, same era.
Gridmm: Grid memory map for vision-and-language navigation
Z. Wang, X. Li, J. Yang, Y. Liu, and S. Jiang · 2023
Cited alongside, same era.
Kerm: Knowledge enhanced reasoning for vision-and-language navigation
X. Li, Z. Wang, J. Yang, Y. Wang, and S. Jiang · 2023
Cited alongside, same era.
Learning navigational visual representations with semantic map supervision
Y. Hong, Y. Zhou, R. Zhang, F. Dernoncourt, T. Bui, S. Gould, and H. Tan · 2023
Cited alongside, same era.
Dreamwalker: Mental planning for continuous vision-language navigation
H. Wang, W. Liang, L. Van Gool, and W. Wang · 2023
Cited alongside, same era.
Navid: Video-based vlm plans the next step for vision-and-language navigation
J. Zhang, K. Wang, R. Xu, G. Zhou, Y. Hong, X. Fang, Q. Wu, Z. Zhang, and H. Wang · 2024
Closest in time.
Etpnav: Evolving topological planning for vision-language navigation in continuous environments
D. An, H. Wang, W. Wang, Z. Wang, Y. Huang, K. He, and L. Wang · 2024
Closest in time.
Lookahead exploration with neural radiance representation for continuous vision-language navigation
Z. Wang, X. Li, J. Yang, Y. Liu, J. Hu, M. Jiang, and S. Jiang · 2024
Closest in time.
Vlfm: Vision-language frontier maps for zero-shot semantic navigation
N. Yokoyama, S. Ha, D. Batra, J. Wang, and B. Bucher · 2024
Closest in time.
Imagine before go: Self-supervised generative map for object goal navigation
S. Zhang, X. Yu, X. Song, X. Wang, and S. Jiang · 2024
Closest in time.
Discuss before moving: Visual language navigation via multi-expert discussions
Y. Long, X. Li, W. Cai, and H. Dong · 2024
Closest in time.
Learning generalizable feature fields for mobile manipulation
R.-Z. Qiu, Y. Hu, G. Yang, Y. Song, Y. Fu, J. Ye, J. Mu, R. Yang, N. Atanasov, S. Scherer, et al · 2024
Closest in time.
Goat-bench: A benchmark for multi-modal lifelong navigation
M. Khanna*, R. Ramrakhya*, G. Chhablani, S. Yenamandra, T. Gervet, M. Chang, Z. Kira, D. S. Chaplot, D. Batra, and R. Mottaghi · 2024
Closest in time.