Fetching the paper…
Reading the bibliography…
Following language instructions to navigate in unseen environments is a challenging task for autonomous embodied agents.
“On evaluation of embodied navigation agents,”
Peter Anderson, Angel X. Chang, Devendra Singh Chaplot, et al., · 2018
Earlier work this paper cites.
“Robust navigation with language pretraining and stochastic sampling,”
Xiujun Li, Chunyuan Li, Qiaolin Xia, et al., · 2019
Earlier work this paper cites.
“Learning to navigate unseen environments: Back translation with environmental dropout,”
Hao Tan, Licheng Yu, et al., · 2019
Earlier work this paper cites.
“Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation,”
Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, et al., · 2019
Earlier work this paper cites.
“Self-monitoring navigation agent via auxiliary progress estimation,”
Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan AlRegib, et al., · 2019
Earlier work this paper cites.
“Tactical rewind: Self-correction via backtracking in vision-and-language navigation,”
Liyiming Ke, Xiujun Li, et al., · 2019
Earlier work this paper cites.
“Reverie: Remote embodied visual referring expression in real indoor environments,”
Yuankai Qi, Qi Wu, Peter Anderson, et al., · 2020
Earlier work this paper cites.
“Improving vision-and-language navigation with image-text pairs from the web,”
Arjun Majumdar, Ayush Shrivastava, et al., · 2020
Earlier work this paper cites.
“Towards learning a generic agent for vision-and-language navigation via pre-training,”
Weituo Hao, Chunyuan Li, Xiujun Li, et al., · 2020
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, et al., · 2020
Cited alongside, same era.
“Towards learning a generic agent for vision-and-language navigation via pre-training,”
Weituo Hao, Chunyuan Li, Xiujun Li, et al., · 2020
Cited alongside, same era.
“Oscar: Object-semantics aligned pre-training for vision-language tasks,”
Xiujun Li, Xi Yin, Li, et al., · 2020
Cited alongside, same era.
“Vision-language navigation with self-supervised auxiliary reasoning tasks,”
Fengda Zhu, Yi Zhu, Xiaojun Chang, and Xiaodan Liang, · 2020
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, et al., · 2021
Cited alongside, same era.
“CPT: colorful prompt tuning for pre-trained vision-language models,”
“Neighbor-view enhanced model for vision and language navigation,”
Dong An, Yuankai Qi, Qi Wu, et al., · 2021
Later among the works it cites.
“Multi-speaker pitch tracking via embodied self-supervised learning,”
Xiang Li, Yifan Sun, Xihong Wu, et al., · 2022
Later among the works it cites.
“ADAPT: vision-language navigation with modality-aligned action prompts,”
Bingqian Lin, Yi Zhu, Zicong Chen, et al., · 2022
Later among the works it cites.
“Visual-language navigation pretraining via prompt-based environmental self-exploration,”
Xiwen Liang, Fengda Zhu, et al., · 2022
Later among the works it cites.
“Reinforced structured state-evolution for vision-language navigation,”
J. Chen, C. Gao, E. Meng, et al., · 2022
Later among the works it cites.
“Ndc-scene: Boost monocular 3d semantic scene completion in normalized device coordinates space,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuan Yao, Ao Zhang, Zhiyuan Liu, et al., · 2021
Cited alongside, same era.
“VLN BERT: A recurrent vision-and-language BERT for navigation,”
Yicong Hong, Qi Wu, et al., · 2021
Cited alongside, same era.
“Vision-language navigation with random environmental mixup,”
Chong Liu, Fengda Zhu, Xiaojun Chang, et al., · 2021
Cited alongside, same era.
“The road to know-where: An object-and-room informed sequential BERT for indoor vision-language navigation,”
Yuankai Qi, Zizheng Pan, Yicong Hong, Qi Wu, et al., · 2021
Cited alongside, same era.
Jiawei Yao, Chuming Li, Keqiang Sun, Yingjie Cai, Hao Li, Wanli Ouyang, and Hongsheng Li, · 2023
Closest in time.
Jiawei Yao and Jusheng Zhang, · 2023
Closest in time.
“VGDiffZero: Text-to-image diffusion models can be zero-shot visual grounders,”
Xuyang Liu, Siteng Huang, Yachen Kang, et al., · 2023
Closest in time.
“Reinforced vision-and-language navigation based on historical BERT,”
Zixuan Zhang, Shuhan Qi, Zihao Zhou, et al., · 2023
Closest in time.