Fetching the paper…
Reading the bibliography…
Vision-and-Language Navigation (VLN) requires the agent to follow language instructions to navigate through 3D environments.
Unrolled generative adversarial networks
L. Metz, B. Poole, D. Pfau, and J. Sohl-Dickstein · 2016
Earlier work this paper cites.
Adversarial nets with perceptual losses for text-to-image synthesis
M. Cha, Y. Gwon, and H. Kung · 2017
Earlier work this paper cites.
Matterport3d: Learning from rgb-d data in indoor environments
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang · 2017
Earlier work this paper cites.
Improved training of wasserstein gans
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville · 2017
Earlier work this paper cites.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. N. Metaxas · 2017
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell · 2018
Earlier work this paper cites.
On the convergence properties of gan training
L. Mescheder · 2018
Earlier work this paper cites.
Stackgan++: Realistic image synthesis with stacked generative adversarial networks
H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. N. Metaxas · 2018
Earlier work this paper cites.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
H. Chen, A. Suhr, D. Misra, N. Snavely, and Y. Artzi · 2019
Earlier work this paper cites.
Stay on the path: Instruction fidelity in vision-and-language navigation
V. Jain, G. Magalhaes, A. Ku, A. Vaswani, E. Ie, and J. Baldridge · 2019
Earlier work this paper cites.
Robust navigation with language pretraining and stochastic sampling
X. Li, C. Li, Q. Xia, Y. Bisk, A. Celikyilmaz, J. Gao, N. Smith, and Y. Choi · 2019
Earlier work this paper cites.
Self-monitoring navigation agent via auxiliary progress estimation
C.-Y. Ma, J. Lu, Z. Wu, G. AlRegib, Z. Kira, R. Socher, and C. Xiong · 2019
Earlier work this paper cites.
K. Nguyen and H. Daumé III · 2019
Earlier work this paper cites.
Semantic image synthesis with spatially-adaptive normalization
T. Park, M.-Y. Liu, T.-C. Wang, and J.-Y. Zhu · 2019
Earlier work this paper cites.
Learning to navigate unseen environments: Back translation with environmental dropout
H. Tan, L. Yu, and M. Bansal · 2019
Earlier work this paper cites.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y.-F. Wang, W. Y. Wang, and L. Zhang · 2019
Earlier work this paper cites.
Very long natural scenery image prediction by outpainting
Z. Yang, J. Dong, P. Liu, Y. Yang, and S. Yan · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Generative adversarial networks
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2020
Cited alongside, same era.
Towards learning a generic agent for vision-and-language navigation via pre-training
W. Hao, C. Li, X. Li, L. Carin, and J. Gao · 2020
Cited alongside, same era.
A recurrent vision-and-language bert for navigation
Y. Hong, Q. Wu, Y. Qi, C. Rodriguez-Opazo, and S. Gould · 2020
Cited alongside, same era.
Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding
A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge · 2020
Cited alongside, same era.
Improving vision-and-language navigation with image-text pairs from the web
A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra · 2020
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Habitat-matterport 3d dataset (HM3d): 1000 large-scale 3d environments for embodied AI
S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. M. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang, M. Savva, Y. Zhao, and D. Batra · 2021
Later among the works it cites.
Structured scene memory for vision-language navigation
H. Wang, W. Wang, W. Liang, C. Xiong, and J. Shen · 2021
Later among the works it cites.
ediffi: Text-to-image diffusion models with an ensemble of expert denoisers
Y. Balaji, S. Nah, X. Huang, A. Vahdat, J. Song, K. Kreis, M. Aittala, T. Aila, S. Laine, B. Catanzaro, et al · 2022
Later among the works it cites.
Learning from unlabeled 3d environments for vision-and-language navigation
S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reverie: Remote embodied visual referring expression in real indoor environments
Y. Qi, Q. Wu, P. Anderson, X. Wang, W. Y. Wang, C. Shen, and A. v. d. Hengel · 2020
Cited alongside, same era.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox · 2020
Cited alongside, same era.
Vision-and-dialog navigation
J. Thomason, M. Murray, M. Cakmak, and L. Zettlemoyer · 2020
Cited alongside, same era.
Environment-agnostic multitask learning for natural language grounded navigation
X. E. Wang, V. Jain, E. Ie, W. Y. Wang, Z. Kozareva, and S. Ravi · 2020
Cited alongside, same era.
mt5: A massively multilingual pre-trained text-to-text transformer
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel · 2020
Cited alongside, same era.
Text-guided neural image inpainting
L. Zhang, Q. Chen, B. Hu, and S. Jiang · 2020
Cited alongside, same era.
History aware multimodal transformer for vision-and-language navigation
S. Chen, P.-L. Guhur, C. Schmid, and I. Laptev · 2021
Cited alongside, same era.
Simple and effective synthesis of indoor 3d scenes
J. Y. Koh, H. Agrawal, D. Batra, R. Tucker, A. Waters, H. Lee, Y. Yang, J. Baldridge, and P. Anderson · 2022
Later among the works it cites.
Teach: Task-driven embodied agents that chat
A. Padmakumar, J. Thomason, A. Shrivastava, P. Lange, A. Narayan-Chen, S. Gella, R. Piramuthu, G. Tur, and D. Hakkani-Tur · 2022
Later among the works it cites.
Hop: history-and-order aware pre-training for vision-and-language navigation
Y. Qiao, Y. Qi, Y. Hong, Z. Yu, P. Wang, and Q. Wu · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Later among the works it cites.
Less is more: Generating grounded navigation instructions from landmarks
S. Wang, C. Montgomery, J. Orbay, V. Birodkar, A. Faust, I. Gur, N. Jaques, A. Waters, J. Baldridge, and P. Anderson · 2022
Later among the works it cites.
Target-driven structured transformer planner for vision-language navigation
Y. Zhao, J. Chen, C. Gao, W. Wang, L. Yang, H. Ren, H. Xia, and S. Liu · 2022
Later among the works it cites.
Multidiffusion: Fusing diffusion paths for controlled image generation
O. Bar-Tal, L. Yariv, Y. Lipman, and T. Dekel · 2023
Closest in time.
Scaling up gans for text-to-image synthesis
M. Kang, J.-Y. Zhu, R. Zhang, J. Park, E. Shechtman, S. Paris, and T. Park · 2023
Closest in time.
Improving vision-and-language navigation by generating future-view image semantics
J. Li and M. Bansal · 2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Closest in time.
Roomdreamer: Text-driven 3d indoor scene synthesis with coherent geometry and texture
L. Song, L. Cao, H. Xu, K. Kang, F. Tang, J. Yuan, and Y. Zhao · 2023
Closest in time.