Fetching the paper…
Reading the bibliography…
Real-world navigation often involves dealing with unexpected obstructions such as closed doors, moved objects, and unpredictable entities.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Real-time obstacle avoidance for fast mobile robots
Johann Borenstein and Yoram Koren. 1989 · 1989
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Distinctive image features from scale-invariant keypoints
David G Lowe. 2004 · 2004
Earlier work this paper cites.
A new model for learning in graph domains. In Proceedings. 2005 IEEE International Joint Conference on Neural Networks, 2005. , Vol. 2. IEEE, 729–734
Marco Gori, Gabriele Monfardini, and Franco Scarselli. 2005 · 2005
Earlier work this paper cites.
Drag-and-drop pasting
Jiaya Jia, Jian Sun, Chi-Keung Tang, and Heung-Yeung Shum. 2006 · 2006
Earlier work this paper cites.
Photo clip art
Jean-François Lalonde, Derek Hoiem, Alexei A Efros, Carsten Rother, John Winn, and Antonio Criminisi. 2007 · 2007
Earlier work this paper cites.
Curriculum learning. In Proceedings of the 26th annual international conference on machine learning . 41–48
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
Gaussian mixture models
Douglas A Reynolds et al · 2009
Earlier work this paper cites.
Planning and obstacle avoidance in mobile robotics
Antonio Sgorbissa and Renato Zaccaria. 2012 · 2012
Earlier work this paper cites.
Automatic scene inference for 3d object compositing
Kevin Karsch, Kalyan Sunkavalli, Sunil Hadap, Nathan Carr, Hailin Jin, Rafael Fonte, Michael Sittig, and David Forsyth. 2014 · 2014
Earlier work this paper cites.
Matterport3D: Learning from RGB-D Data in Indoor Environments. In International Conference on 3D Vision (3DV) . 667–676
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niebner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. 2017 · 2017
Earlier work this paper cites.
Mobile robot navigation and obstacle avoidance techniques: A review
Anish Pandey, Shalini Pandey, and DR Parhi. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
On evaluation of embodied navigation agents
Peter Anderson, Angel Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, et al · 2018
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. 2018 · 2018
Earlier work this paper cites.
Context-aware synthesis and placement of object instances
Donghoon Lee, Sifei Liu, Jinwei Gu, Ming-Yu Liu, Ming-Hsuan Yang, and Jan Kautz. 2018 · 2018
Earlier work this paper cites.
Gibson env: Real-world perception for embodied agents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9068–9079
Fei Xia, Amir R Zamir, Zhiyang He, Alexander Sax, Jitendra Malik, and Silvio Savarese. 2018 · 2018
Earlier work this paper cites.
Visualizing and Understanding Generative Adversarial Networks. In International Conference on Learning Representations
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Bolei Zhou, Joshua B. Tenenbaum, William T. Freeman, and Antonio Torralba. 2019 · 2019
Earlier work this paper cites.
General evaluation for instruction conditioned navigation using dynamic time warping
Gabriel Ilharco, Vihan Jain, Alexander Ku, Eugene Ie, and Jason Baldridge. 2019 · 2019
Earlier work this paper cites.
Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . 1862–1872
Vihan Jain, Gabriel Magalhaes, Alexander Ku, Ashish Vaswani, Eugene Ie, and Jason Baldridge. 2019 · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization. In International Conference on Learning Representations
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Earlier work this paper cites.
Self-Monitoring Navigation Agent via Auxiliary Progress Estimation. In International Conference on Learning Representations
Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan AlRegib, Zsolt Kira, Richard Socher, and Caiming Xiong. 2019 · 2019
Earlier work this paper cites.
Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2337–2346
Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. 2019 · 2019
Cited alongside, same era.
Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . 2610–2621
Hao Tan, Licheng Yu, and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
Towards learning a generic agent for vision-and-language navigation via pre-training. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 13137–13146
Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, and Jianfeng Gao. 2020 · 2020
Cited alongside, same era.
Language and visual entity relationship graph for agent navigation
Yicong Hong, Cristian Rodriguez, Yuankai Qi, Qi Wu, and Stephen Gould. 2020 · 2020
Cited alongside, same era.
Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d
Yiyi Liao, Jun Xie, and Andreas Geiger. 2022 · 2022
Later among the works it cites.
Multimodal transformer with variable-length memory for vision-and-language navigation. In European Conference on Computer Vision . Springer, 380–397
Chuang Lin, Yi Jiang, Jianfei Cai, Lizhen Qu, Gholamreza Haffari, and Zehuan Yuan. 2022 · 2022
Later among the works it cites.
Object insertion based data augmentation for semantic segmentation. In 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 359–365
Yuan Ren, Siyan Zhao, and Liu Bingbing. 2022 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Later among the works it cites.
Dynamic obstacle avoidance and path planning through reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beyond the nav-graph: Vision-and-language navigation in continuous environments. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVIII 16 . Springer, 104–120
Jacob Krantz, Erik Wijmans, Arjun Majumdar, Dhruv Batra, and Stefan Lee. 2020 · 2020
Cited alongside, same era.
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 4392–4412
Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge. 2020 · 2020
Cited alongside, same era.
Improving vision-and-language navigation with image-text pairs from the web. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16 . Springer, 259–274
Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson, Devi Parikh, and Dhruv Batra. 2020 · 2020
Cited alongside, same era.
Counterfactual vision-and-language navigation: Unravelling the unseen
Amin Parvaneh, Ehsan Abbasnejad, Damien Teney, Javen Qinfeng Shi, and Anton Van den Hengel. 2020 · 2020
Cited alongside, same era.
Reverie: Remote embodied visual referring expression in real indoor environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9982–9991
Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. 2020 · 2020
Cited alongside, same era.
Vision-and-dialog navigation. In Conference on Robot Learning . PMLR, 394–406
Jesse Thomason, Michael Murray, Maya Cakmak, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Vision-language navigation with self-supervised auxiliary reasoning tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10012–10022
Fengda Zhu, Yi Zhu, Xiaojun Chang, and Xiaodan Liang. 2020 · 2020
Cited alongside, same era.
Neighbor-view enhanced model for vision and language navigation. In Proceedings of the 29th ACM International Conference on Multimedia . 5101–5109
Dong An, Yuankai Qi, Yan Huang, Qi Wu, Liang Wang, and Tieniu Tan. 2021 · 2021
Cited alongside, same era.
Khawla Almazrouei, Ibrahim Kamel, and Tamer Rabie. 2023 · 2023
Later among the works it cites.
Grounded entity-landmark adaptive pre-training for vision-and-language navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 12043–12053
Yibo Cui, Liang Xie, Yakun Zhang, Meishan Zhang, Ye Yan, and Erwei Yin. 2023 · 2023
Later among the works it cites.
Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13142–13153
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. 2023 · 2023
Later among the works it cites.
A new path: Scaling vision-and-language navigation with synthetic instructions and imitation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10813–10823
Aishwarya Kamath, Peter Anderson, Su Wang, Jing Yu Koh, Alexander Ku, Austin Waters, Yinfei Yang, Jason Baldridge, and Zarana Parekh. 2023 · 2023
Later among the works it cites.
Renderable neural radiance map for visual navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9099–9108
Obin Kwon, Jeongho Park, and Songhwai Oh. 2023 · 2023
Later among the works it cites.
Improving vision-and-language navigation by generating future-view image semantics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10803–10812
Jialu Li and Mohit Bansal. 2023 · 2023
Later among the works it cites.
Learning vision-and-language navigation from youtube videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 8317–8326
Kunyang Lin, Peihao Chen, Diwei Huang, Thomas H Li, Mingkui Tan, and Chuang Gan. 2023 · 2023
Later among the works it cites.
Aerialvln: Vision-and-language navigation for uavs. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 15384–15394
Shubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang, Yanning Zhang, and Qi Wu. 2023 · 2023
Later among the works it cites.
VLN-PETL: Parameter-Efficient Transfer Learning for Vision-and-Language Navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 15443–15452
Yanyuan Qiao, Zheng Yu, and Qi Wu. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Vision and Language Navigation in the Real World via Online Visual Language Mapping. In 2nd Workshop on Language and Robot Learning: Language as Grounding
Chengguang Xu, Hieu Trung Nguyen, Christopher Amato, and Lawson Wong. 2023 · 2023
Later among the works it cites.
Mind the Gap: Improving Success Rate of Vision-and-Language Navigation by Revisiting Oracle Success Routes. In Proceedings of the 31st ACM International Conference on Multimedia . 4349–4358
Chongyang Zhao, Yuankai Qi, and Qi Wu. 2023 · 2023
Later among the works it cites.
Etpnav: Evolving topological planning for vision-language navigation in continuous environments
Dong An, Hanqing Wang, Wenguan Wang, Zun Wang, Yan Huang, Keji He, and Liang Wang. 2024 · 2024
Closest in time.
3d copy-paste: Physically plausible object insertion for monocular 3d detection
Yunhao Ge, Hong-Xing Yu, Cheng Zhao, Yuliang Guo, Xinyu Huang, Liu Ren, Laurent Itti, and Jiajun Wu. 2024 · 2024
Closest in time.
Panogen: Text-conditioned panoramic environment generation for vision-and-language navigation
Jialu Li and Mohit Bansal. 2024 · 2024
Closest in time.
Focaldreamer: Text-driven 3d editing via focal-fusion assembly. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 3279–3287
Yuhan Li, Yishun Dou, Yue Shi, Yu Lei, Xuanhong Chen, Yi Zhang, Peng Zhou, and Bingbing Ni. 2024 · 2024
Closest in time.
Language-driven Object Fusion into Neural Radiance Fields with Pose-Conditioned Dataset Updates. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5176–5187
Ka Chun Shum, Jaeyeon Kim, Binh-Son Hua, Duc Thanh Nguyen, and Sai-Kit Yeung. 2024 · 2024
Closest in time.
Safe-VLN: Collision Avoidance for Vision-and-Language Navigation of Autonomous Robots Operating in Continuous Environments
Lu Yue, Dongliang Zhou, Liang Xie, Feitian Zhang, Ye Yan, and Erwei Yin. 2024 · 2024
Closest in time.