Fetching the paper…
Reading the bibliography…
Trained with an unprecedented scale of data, large language models (LLMs) like ChatGPT and GPT-4 exhibit the emergence of significant reasoning abilities from model scaling.
Fast r-cnn
R. Girshick · 2015
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel · 2018
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell · 2018
Earlier work this paper cites.
Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation
X. Wang, W. Xiong, H. Wang, and W. Y. Wang · 2018
Earlier work this paper cites.
Self-monitoring navigation agent via auxiliary progress estimation
C.-Y. Ma, J. Lu, Z. Wu, G. AlRegib, Z. Kira, R. Socher, and C. Xiong · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Earlier work this paper cites.
Learning to navigate unseen environments: Back translation with environmental dropout
H. Tan, L. Yu, and M. Bansal · 2019
Earlier work this paper cites.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y.-F. Wang, W. Y. Wang, and L. Zhang · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Evolving graphical planner: Contextual global planning for vision-and-language navigation
Z. Deng, K. Narasimhan, and O. Russakovsky · 2020
Earlier work this paper cites.
Counterfactual vision-and-language navigation via adversarial path sampler
T.-J. Fu, X. E. Wang, M. F. Peterson, S. T. Grafton, M. P. Eckstein, and W. Y. Wang · 2020
Earlier work this paper cites.
Towards learning a generic agent for vision-and-language navigation via pre-training
W. Hao, C. Li, X. Li, L. Carin, and J. Gao · 2020
Earlier work this paper cites.
Language and visual entity relationship graph for agent navigation
Y. Hong, C. Rodriguez, Y. Qi, Q. Wu, and S. Gould · 2020
Earlier work this paper cites.
Sub-instruction aware vision-and-language navigation
Y. Hong, C. Rodriguez-Opazo, Q. Wu, and S. Gould · 2020
Earlier work this paper cites.
Beyond the nav-graph: Vision-and-language navigation in continuous environments
J. Krantz, E. Wijmans, A. Majumdar, D. Batra, and S. Lee · 2020
Earlier work this paper cites.
Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge · 2020
Earlier work this paper cites.
Improving vision-and-language navigation with image-text pairs from the web
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra · 2020
Earlier work this paper cites.
Counterfactual vision-and-language navigation: Unravelling the unseen
A. Parvaneh, E. Abbasnejad, D. Teney, J. Q. Shi, and A. van den Hengel · 2020
Earlier work this paper cites.
Object-and-action aware model for visual language navigation
Y. Qi, Z. Pan, S. Zhang, A. v. d. Hengel, and Q. Wu · 2020
Earlier work this paper cites.
Reverie: Remote embodied visual referring expression in real indoor environments
Y. Qi, Q. Wu, P. Anderson, X. Wang, W. Y. Wang, C. Shen, and A. v. d. Hengel · 2020
Earlier work this paper cites.
Vision-language navigation with self-supervised auxiliary reasoning tasks
F. Zhu, Y. Zhu, X. Chang, and X. Liang · 2020
Earlier work this paper cites.
BabyWalk: Going farther in vision-and-language navigation by taking baby steps
W. Zhu, H. Hu, J. Chen, Z. Deng, V. Jain, E. Ie, and F. Sha · 2020
Earlier work this paper cites.
Topological planning with transformers for vision-and-language navigation
K. Chen, J. K. Chen, J. Chuang, M. Vázquez, and S. Savarese · 2021
Earlier work this paper cites.
History aware multimodal transformer for vision-and-language navigation
S. Chen, P.-L. Guhur, C. Schmid, and I. Laptev · 2021
Earlier work this paper cites.
Room-and-object aware knowledge reasoning for remote embodied referring expression
C. Gao, J. Chen, S. Liu, L. Wang, Q. Zhang, and Q. Wu · 2021
Cited alongside, same era.
Airbert: In-domain pretraining for vision-and-language navigation
P.-L. Guhur, M. Tapaswi, S. Chen, I. Laptev, and C. Schmid · 2021
Cited alongside, same era.
Landmark-rxr: Solving vision-and-language navigation with fine-grained alignment supervision
K. He, Y. Huang, Q. Wu, J. Yang, D. An, S. Sima, and L. Wang · 2021
Cited alongside, same era.
Vision-language navigation with random environmental mixup
C. Liu, F. Zhu, X. Chang, X. Liang, Z. Ge, and Y.-D. Shen · 2021
Cited alongside, same era.
Episodic transformer for vision-and-language navigation
A. Pashevich, C. Schmid, and C. Sun · 2021
Cited alongside, same era.
The road to know-where: An object-and-room informed sequential bert for indoor vision-language navigation
Less is more: Generating grounded navigation instructions from landmarks
S. Wang, C. Montgomery, J. Orbay, V. Birodkar, A. Faust, I. Gur, N. Jaques, A. Waters, J. Baldridge, and P. Anderson · 2022
Later among the works it cites.
Emergent abilities of large language models
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, et al · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Later among the works it cites.
Cross-modal semantic alignment pre-training for vision-and-language navigation
S. Wu, X. Fu, F. Wu, and Z.-J. Zha · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Qi, Z. Pan, Y. Hong, M.-H. Yang, A. van den Hengel, and Q. Wu · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Structured scene memory for vision-language navigation
H. Wang, W. Wang, W. Liang, C. Xiong, and J. Shen · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2021
Cited alongside, same era.
Do as i can and not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, M. Yan, and A. Zeng · 2022
Cited alongside, same era.
Think global, act local: Dual-scale graph transformer for vision-and-language navigation
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Cited alongside, same era.
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Later among the works it cites.
Diagnosing vision-and-language navigation: What really matters
W. Zhu, Y. Qi, P. Narayana, K. Sone, S. Basu, X. E. Wang, Q. Wu, M. P. Eckstein, and W. Y. Wang · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Closest in time.
Can an embodied agent find your" cat-shaped mug"? llm-based zero-shot object navigation
V. S. Dorbala, J. F. Mullen Jr, and D. Manocha · 2023
Closest in time.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Closest in time.
Layout-aware dreamer for embodied referring expression grounding
M. Li, Z. Wang, T. Tuytelaars, and M.-F. Moens · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
B. Peng, M. Galley, P. He, H. Cheng, Y. Xie, Y. Hu, Q. Huang, L. Liden, Z. Yu, W. Chen, et al · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Closest in time.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
D. Shah, B. Osiński, S. Levine, et al · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface, 2023
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Chatgpt for robotics: Design principles and model abilities
S. Vemprala, R. Bonatti, A. Bucker, and A. Kapoor · 2023
Closest in time.
Visual chatgpt: Talking, drawing and editing with visual foundation models
C. Wu, S. Yin, W. Qi, X. Wang, Z. Tang, and N. Duan · 2023
Closest in time.
Esc: Exploration with soft commonsense constraints for zero-shot object navigation
K. Zhou, K. Zheng, C. Pryor, Y. Shen, H. Jin, L. Getoor, and X. E. Wang · 2023
Closest in time.
Chatgpt asks, blip-2 answers: Automatic questioning towards enriched visual descriptions, 2023
D. Zhu, J. Chen, K. Haydarov, X. Shen, W. Zhang, and M. Elhoseiny · 2023
Closest in time.