Fetching the paper…
Reading the bibliography…
Top-down images play an important role in safety-critical settings such as autonomous navigation and aerial surveillance, where they provide holistic spatial information that front-view images cannot capture.
Microsoft coco: Common objects in context, 2015
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge, 2015
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
CARLA: An open urban driving simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun · 2017
Earlier work this paper cites.
A survey on vision-based uav navigation
Yuncheng Lu, Zhucun Xue, Gui-Song Xia, and Liangpei Zhang · 2017
Earlier work this paper cites.
Semantic segmentation drone dataset, 2019
ICG · 2019
Earlier work this paper cites.
An ensemble deep learning method with optimized weights for drone-based water rescue and surveillance
Jan Gasienica-Jozkowy, Mateusz Knapik, and Boguslaw Cyganek · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Detection and tracking meet drones challenge
Pengfei Zhu, Longyin Wen, Dawei Du, Xiao Bian, Heng Fan, Qinghua Hu, and Haibin Ling · 2021
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models, 2022
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev · 2022
Earlier work this paper cites.
Seadronessee: A maritime benchmark for detecting humans in open water
Leon Amadeus Varga, Benjamin Kiefer, Martin Messmer, and Andreas Zell · 2022
Earlier work this paper cites.
Rsgpt: A remote sensing vision language model and benchmark, 2023
Yuan Hu, Jianlong Yuan, Congcong Wen, Xiaonan Lu, and Xiang Li · 2023
Earlier work this paper cites.
Geochat: Grounded large vision-language model for remote sensing, 2023
Kartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer, Abhijit Das, Salman Khan, and Fahad Shahbaz Khan · 2023
Earlier work this paper cites.
Evaluating object hallucination in large vision-language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen · 2023
Earlier work this paper cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Earlier work this paper cites.
Sigmoid loss for language image pre-training, 2023
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Cited alongside, same era.
Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Haodong Duan, Junming Yang, Yuxuan Qiao, Xinyu Fang, Lin Chen, Yuan Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Jiaqi Wang, et al · 2024
Cited alongside, same era.
Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, and Rongrong Ji · 2024
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
Google · 2024
Cited alongside, same era.
Claude 4.1, 2025b
Anthropic · 2025
Closest in time.
Claude 4, 2025c
Anthropic · 2025
Closest in time.
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, et al · 2025
Closest in time.
Geobench-vlm: Benchmarking vision-language models for geospatial tasks, 2025
Muhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Kartik Kuckreja, Fahad Shahbaz Khan, Paolo Fraccaro, Alexandre Lacoste, and Salman Khan · 2025
Closest in time.
Kimi-vl technical report, 2025
KimiTeam · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou · 2024
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao · 2024
Cited alongside, same era.
Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model, 2024
Dilxat Muhtar, Zhenshi Li, Feng Gu, Xueliang Zhang, and Pengfeng Xiao · 2024
Cited alongside, same era.
OpenAI · 2024
Cited alongside, same era.
Mllm can see? dynamic correction decoding for hallucination mitigation, 2024
Chenxi Wang, Xiang Chen, Ningyu Zhang, Bozhong Tian, Haoming Xu, Shumin Deng, and Huajun Chen · 2024
Cited alongside, same era.
Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024
Zhiyu Wu, Xiaokang Chen, Zizheng Pan, et al · 2024
Cited alongside, same era.
Mm-vet: Evaluating large multimodal models for integrated capabilities, 2024
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang · 2024
Cited alongside, same era.
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen · 2024
Cited alongside, same era.
Nearmap · 2025
Closest in time.
Introducing gpt-4.1 in the api, 2025a
OpenAI · 2025
Closest in time.
Introducing gpt-5, 2025b
OpenAI · 2025
Closest in time.
Introducing openai o3 and o4-mini, 2025c
OpenAI · 2025
Closest in time.
Aerial traffic images
Shaha · 2025
Closest in time.
Vlm-r1: A stable and generalizable r1-style large vision-language model, 2025
Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, Ruochen Xu, and Tiancheng Zhao · 2025
Closest in time.
Internvl3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency, 2025
Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, et al · 2025
Closest in time.
FlexiFly: Interfacing the Physical World with Foundation Models Empowered by Reconfigurable Drone Systems , pp. 463–476
Minghui Zhao, Junxi Xia, Kaiyuan Hou, Yanchen Liu, Stephen Xia, and Xiaofan Jiang · 2025
Closest in time.
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025
Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, et al · 2025
Closest in time.