Fetching the paper…
Reading the bibliography…
Aerial Visual Object Search (AVOS) tasks in urban environments require Unmanned Aerial Vehicles (UAVs) to autonomously search for and identify target objects using visual and textual cues without external guidance.
Matterport3d: Learning from rgb-d data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. 2017 · 2017
Earlier work this paper cites.
DBSCAN revisited, revisited: why and how you should (still) use DBSCAN
Erich Schubert, Jörg Sander, Martin Ester, Hans Peter Kriegel, and Xiaowei Xu. 2017 · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3674–3683
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel. 2018 · 2018
Earlier work this paper cites.
Airsim: High-fidelity visual and physical simulation for autonomous vehicles. In Field and Service Robotics: Results of the 11th International Conference . Springer, 621–635
Shital Shah, Debadeepta Dey, Chris Lovett, and Ashish Kapoor. 2018 · 2018
Earlier work this paper cites.
Gibson env: Real-world perception for embodied agents. In Proceedings of the IEEE conference on computer vision and pattern recognition . 9068–9079
Fei Xia, Amir R Zamir, Zhiyang He, Alexander Sax, Jitendra Malik, and Silvio Savarese. 2018 · 2018
Earlier work this paper cites.
Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames
Erik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee, Irfan Essa, Devi Parikh, Manolis Savva, and Dhruv Batra. 2019 · 2019
Earlier work this paper cites.
UAV autonomous target search based on deep reinforcement learning in complex disaster scene
Chunxue Wu, Bobo Ju, Yan Wu, Xiao Lin, Naixue Xiong, Guangquan Xu, Hongyan Li, and Xuefeng Liang. 2019 · 2019
Earlier work this paper cites.
UAV-assisted emergency networks in disasters
Nan Zhao, Weidang Lu, Min Sheng, Yunfei Chen, Jie Tang, F Richard Yu, and Kai-Kit Wong. 2019 · 2019
Earlier work this paper cites.
Object goal navigation using goal-oriented semantic exploration
Devendra Singh Chaplot, Dhiraj Prakashchand Gandhi, Abhinav Gupta, and Russ R Salakhutdinov. 2020 · 2020
Earlier work this paper cites.
Indoor navigation for mobile agents: A multimodal vision fusion model. In 2020 international joint conference on neural networks (IJCNN) . IEEE, 1–8
Dongfang Liu, Yiming Cui, Zhiwen Cao, and Yingjie Chen. 2020 · 2020
Earlier work this paper cites.
Reverie: Remote embodied visual referring expression in real indoor environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9982–9991
Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. 2020 · 2020
Earlier work this paper cites.
Active Object Search. In Proceedings of the 28th ACM International Conference on Multimedia . 973–981
Jie Wu, Tianshui Chen, Lishan Huang, Hefeng Wu, Guanbin Li, Ling Tian, and Liang Lin. 2020 · 2020
Earlier work this paper cites.
Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai
Santhosh K Ramakrishnan, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alex Clegg, John Turner, Eric Undersander, Wojciech Galuba, Andrew Westbury, Angel X Chang, et al · 2021
Earlier work this paper cites.
Efficiency of UAV-based last-mile delivery under congestion in low-altitude air
Ruifeng She and Yanfeng Ouyang. 2021 · 2021
Earlier work this paper cites.
Auxiliary tasks and exploration enable objectgoal navigation. In Proceedings of the IEEE/CVF international conference on computer vision . 16117–16126
Joel Ye, Dhruv Batra, Abhishek Das, and Erik Wijmans. 2021 · 2021
Earlier work this paper cites.
ProcTHOR: Large-Scale Embodied AI Using Procedural Generation
Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Kiana Ehsani, Jordi Salvador, Winson Han, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. 2022 · 2022
Earlier work this paper cites.
Poni: Potential functions for objectgoal navigation with interaction-free learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18890–18900
Santhosh Kumar Ramakrishnan, Devendra Singh Chaplot, Ziad Al-Halah, Jitendra Malik, and Kristen Grauman. 2022 · 2022
Cited alongside, same era.
Multi-UAV cooperative system for search and rescue based on YOLOv5
Linjie Xing, Xiaoyan Fan, Yaxin Dong, Zenghui Xiong, Lin Xing, Yang Yang, Haicheng Bai, and Chengjiang Zhou. 2022 · 2022
Cited alongside, same era.
A deep reinforcement learning based searching method for source localization
Yong Zhao, Bin Chen, XiangHan Wang, Zhengqiu Zhu, Yiduo Wang, Guangquan Cheng, Rui Wang, Rongxiao Wang, Ming He, and Yu Liu. 2022 · 2022
Cited alongside, same era.
Object goal navigation with recursive implicit maps. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 7089–7096
Shizhe Chen, Thomas Chabal, Ivan Laptev, and Cordelia Schmid. 2023 · 2023
Cited alongside, same era.
Can an embodied agent find your “cat-shaped mug”? llm-based zero-shot object navigation
Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 5228–5234
Wenzhe Cai, Siyuan Huang, Guangran Cheng, Yuxing Long, Peng Gao, Changyin Sun, and Hao Dong. 2024b · 2024
Later among the works it cites.
Zhixi Cai, Cristian Rojas Cardenas, Kevin Leo, Chenyuan Zhang, Kal Backman, Hanbing Li, Boying Li, Mahsa Ghorbanali, Stavya Datta, Lizhen Qu, et al · 2024
Later among the works it cites.
Say-REAPEx: An LLM-Modulo UAV Online Planning Framework for Search and Rescue. In 2nd CoRL Workshop on Learning Effective Abstractions for Planning
Björn Döschl and Jane Jean Kiam. 2024 · 2024
Later among the works it cites.
EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment
Chen Gao, Baining Zhao, Weichen Zhang, Jinzhu Mao, Jun Zhang, Zhiheng Zheng, Fanhang Man, Jianjie Fang, Zile Zhou, Jinqiang Cui, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vishnu Sashank Dorbala, James F Mullen, and Dinesh Manocha. 2023 · 2023
Cited alongside, same era.
Object-goal visual navigation via effective exploration of relations among historical navigation states. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2563–2573
Heming Du, Lincheng Li, Zi Huang, and Xin Yu. 2023 · 2023
Cited alongside, same era.
Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 23171–23181
Samir Yitzhak Gadre, Mitchell Wortsman, Gabriel Ilharco, Ludwig Schmidt, and Shuran Song. 2023 · 2023
Cited alongside, same era.
UAV swarm cooperative target search: A multi-agent reinforcement learning approach
Yukai Hou, Jin Zhao, Rongqing Zhang, Xiang Cheng, and Liuqing Yang. 2023 · 2023
Cited alongside, same era.
Aerialvln: Vision-and-language navigation for uavs. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 15384–15394
Shubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang, Yanning Zhang, and Qi Wu. 2023 · 2023
Cited alongside, same era.
Co-evolutionary algorithm-based multi-unmanned aerial vehicle cooperative path planning
Yan Wu, Mingtao Nie, Xiaolei Ma, Yicong Guo, and Xiaoxiong Liu. 2023 · 2023
Cited alongside, same era.
Habitat-matterport 3d semantics dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4927–4936
Karmesh Yadav, Ram Ramrakhya, Santhosh Kumar Ramakrishnan, Theo Gervet, John Turner, Aaron Gokaslan, Noah Maestre, Angel Xuan Chang, Dhruv Batra, Manolis Savva, et al · 2023
Cited alongside, same era.
Co-navgpt: Multi-robot cooperative visual semantic navigation using large language models
Bangguo Yu, Hamidreza Kasaei, and Ming Cao. 2023a · 2023
Cited alongside, same era.
Later among the works it cites.
Aerial Vision-and-Language Navigation via Semantic-Topo-Metric Representation Guided LLM Reasoning
Yunpeng Gao, Zhigang Wang, Linglin Jing, Dong Wang, Xuelong Li, and Bin Zhao. 2024a · 2024
Later among the works it cites.
CityNav: Language-Goal Aerial Navigation Dataset with Geographic Information
Jungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Yutaka Matsuo, and Nakamasa Inoue. 2024 · 2024
Later among the works it cites.
Aligning cyber space with physical world: A comprehensive survey on embodied ai
Yang Liu, Weixing Chen, Yongjie Bai, Xiaodan Liang, Guanbin Li, Wen Gao, and Liang Lin. 2024 · 2024
Later among the works it cites.
Towards realistic uav vision-language navigation: Platform, benchmark, and methodology
Xiangyu Wang, Donglin Yang, Ziqin Wang, Hohin Kwan, Jinyu Chen, Wenjun Wu, Hongsheng Li, Yue Liao, and Si Liu. 2024 · 2024
Later among the works it cites.
Voronav: Voronoi-based zero-shot object navigation with large language model
Pengying Wu, Yao Mu, Bingxian Wu, Yi Hou, Ji Ma, Shanghang Zhang, and Chang Liu. 2024 · 2024
Later among the works it cites.
Fanglong Yao, Yuanchang Yue, Youzhi Liu, Xian Sun, and Kun Fu. 2024 · 2024
Later among the works it cites.
Navgpt: Explicit reasoning in vision-and-language navigation with large language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 7641–7649
Gengze Zhou, Yicong Hong, and Qi Wu. 2024 · 2024
Later among the works it cites.
OpenFly: A Versatile Toolchain and Large-scale Benchmark for Aerial Vision-Language Navigation
Yunpeng Gao, Chenhui Li, Zhongrui You, Junli Liu, Zhen Li, Pengan Chen, Qizhi Chen, Zhonghan Tang, Liansheng Wang, Penghui Yang, et al · 2025
Closest in time.
WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation
Dujun Nie, Xianda Guo, Yiqun Duan, Ruijun Zhang, and Long Chen. 2025 · 2025
Closest in time.
StrucGCN: Structural enhanced graph convolutional networks for graph embedding
Jie Zhang, Mingxuan Li, Yitai Xu, Hua He, Qun Li, and Tao Wang. 2025 · 2025
Closest in time.
Cityeqa: A hierarchical llm agent on embodied question answering benchmark in city space
Yong Zhao, Kai Xu, Zhengqiu Zhu, Yue Hu, Zhiheng Zheng, Yingfeng Chen, Yatai Ji, Chen Gao, Yong Li, and Jincai Huang. 2025 · 2025
Closest in time.