Fetching the paper…
Reading the bibliography…
Visual language navigation (VLN) is an embodied task demanding a wide range of skills encompassing understanding, perception, and planning.
1907
Earlier work this paper cites.
1912
Earlier work this paper cites.
H. B. Kal and H. Struikmans, “Breast carcinoma during pregnancy: International recommendations from an expert meeting,” Cancer , vol. 107, no. 4, 2006
2006
Earlier work this paper cites.
W. H. Organization et al. , “Who expert consultation on public health intervention against early childhood caries: report of a meeting, bangkok, thailand, 26-28 january 2016,” World Health Organization, Tech. Rep., 2017
2017
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 3674–3683
2018
Earlier work this paper cites.
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell, “Speaker-follower models for vision-and-language navigation,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Earlier work this paper cites.
J. Stojkov, G. Bowers, M. Draper, T. Duffield, P. Duivenvoorden, M. Groleau, D. Haupstein, R. Peters, J. Pritchard, C. Radom et al. , “Hot topic: Management of cull dairy cows—consensus of an expert consultation in canada,” Journal of dairy science , vol. 101, no. 12, pp. 11 170–11 174, 2018
2018
Earlier work this paper cites.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. van den Hengel, “Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Earlier work this paper cites.
H. Tan, L. Yu, and M. Bansal, “Learning to navigate unseen environments: Back translation with environmental dropout,” in Proceedings of NAACL-HLT , 2019, pp. 2610–2621
2019
Earlier work this paper cites.
H. Tan and M. Bansal, “Lxmert: Learning cross-modality encoder representations from transformers,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 5100–5111
2019
Earlier work this paper cites.
J. Krantz, E. Wijmans, A. Majumdar, D. Batra, and S. Lee, “Beyond the nav-graph: Vision-and-language navigation in continuous environments,” in European Conference on Computer Vision . Springer, 2020, pp. 104–120
2020
Earlier work this paper cites.
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge, “Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 4392–4412
2020
Earlier work this paper cites.
W. Hao, C. Li, X. Li, L. Carin, and J. Gao, “Towards learning a generic agent for vision-and-language navigation via pre-training,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 13 137–13 146
2020
Earlier work this paper cites.
A. Parvaneh, E. Abbasnejad, D. Teney, J. Q. Shi, and A. van den Hengel, “Counterfactual vision-and-language navigation: Unravelling the unseen,” Advances in Neural Information Processing Systems , vol. 33, pp. 5296–5307, 2020
2020
Earlier work this paper cites.
T.-J. Fu, X. E. Wang, M. F. Peterson, S. T. Grafton, M. P. Eckstein, and W. Y. Wang, “Counterfactual vision-and-language navigation via adversarial path sampler,” in European Conference on Computer Vision . Springer, 2020, pp. 71–86
2020
Earlier work this paper cites.
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra, “Improving vision-and-language navigation with image-text pairs from the web,” in European Conference on Computer Vision . Springer, 2020, pp. 259–274
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
K. He, Y. Huang, Q. Wu, J. Yang, D. An, S. Sima, and L. Wang, “Landmark-rxr: Solving vision-and-language navigation with fine-grained alignment supervision,” Advances in Neural Information Processing Systems , vol. 34, pp. 652–663, 2021
2021
Cited alongside, same era.
Y. Hong, Q. Wu, Y. Qi, C. Rodriguez-Opazo, and S. Gould, “VLN ↻ \circlearrowright BERT: A recurrent vision-and-language bert for navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 1643–1653
2021
Cited alongside, same era.
S. Chen, P.-L. Guhur, C. Schmid, and I. Laptev, “History aware multimodal transformer for vision-and-language navigation,” Advances in Neural Information Processing Systems , vol. 34, pp. 5834–5847, 2021
2021
S. Wu, X. Fu, F. Wu, and Z.-J. Zha, “Cross-modal semantic alignment pre-training for vision-and-language navigation,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 4233–4241
2022
Later among the works it cites.
S. Y. Gadre, M. Wortsman, G. Ilharco, L. Schmidt, and S. Song, “Clip on wheels: Open-vocabulary models are (almost) zero-shot object navigators,” arXiv , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
OpenAI, “Introducing chatgpt,” https://www.openai.com/blog/chatgpt , 2022
2022
Later among the works it cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in Neural Information Processing Systems , vol. 35, pp. 24 824–24 837, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
W. H. Organization et al. , “Meeting report of the who expert consultation on the definition of extensively drug-resistant tuberculosis, 27-29 october 2020,” 2021
2021
Cited alongside, same era.
C. Liu, F. Zhu, X. Chang, X. Liang, Z. Ge, and Y.-D. Shen, “Vision-language navigation with random environmental mixup,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1644–1654
2021
Cited alongside, same era.
H. Wang, W. Wang, W. Liang, C. Xiong, and J. Shen, “Structured scene memory for vision-language navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 8455–8464
2021
Cited alongside, same era.
A. Pashevich, C. Schmid, and C. Sun, “Episodic transformer for vision-and-language navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 15 942–15 952
2021
Cited alongside, same era.
P.-L. Guhur, M. Tapaswi, S. Chen, I. Laptev, and C. Schmid, “Airbert: In-domain pretraining for vision-and-language navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1634–1643
2021
Cited alongside, same era.
Y. Qi, Z. Pan, Y. Hong, M.-H. Yang, A. van den Hengel, and Q. Wu, “The road to know-where: An object-and-room informed sequential bert for indoor vision-language navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1655–1664
2021
Cited alongside, same era.
Y. Zhu, Y. Weng, F. Zhu, X. Liang, Q. Ye, Y. Lu, and J. Jiao, “Self-motivated communication agent for real-world vision-dialog navigation,” in 2021 IEEE/CVF International Conference on Computer Vision (ICCV) , 2021, pp. 1574–1583
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang, “Glm: General language model pretraining with autoregressive blank infilling,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 320–335
2022
Later among the works it cites.
2023
Closest in time.
D. Shah, B. Osiński, S. Levine et al. , “Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,” in Conference on Robot Learning . PMLR, 2023, pp. 492–504
2023
Closest in time.
——, “Gpt-4 technical report,” 2023
2023
Closest in time.
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” 2023
2023
Closest in time.
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, “Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface,” 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.