Fetching the paper…
Reading the bibliography…
Geolocation, the task of identifying an image's location, requires complex reasoning and is crucial for navigation, monitoring, and cultural preservation.
The StreetLearn environment and dataset
Piotr Mirowski, Andras Banki-Horvath, Keith Anderson, Denis Teplyashin, Karl Moritz Hermann, Mateusz Malinowski, Matthew Koichi Grimes, Karen Simonyan, Koray Kavukcuoglu, Andrew Zisserman, and 1 others. 2019 · 1903
Earlier work this paper cites.
Evaluating cross-modal generative models using retrieval task
Shivangi Bithel and Srikanta Bedathur. 2023 · 1965
Earlier work this paper cites.
IM2GPS: estimating geographic information from a single image
James Hays and Alexei A Efros. 2008 · 2008
Earlier work this paper cites.
Image geo-localization based on multiplenearest neighbor feature matching usinggeneralized graphs
Amir Roshan Zamir and Mubarak Shah. 2014 · 2014
Earlier work this paper cites.
Wide-area image geolocalization with aerial reference imagery
Scott Workman, Richard Souvenir, and Nathan Jacobs. 2015 · 2015
Earlier work this paper cites.
The benchmarking initiative for multimedia evaluation: Mediaeval 2016
Martha Larson, Mohammad Soleymani, Guillaume Gravier, Bogdan Ionescu, and Gareth JF Jones. 2017 · 2016
Earlier work this paper cites.
Exploring the choice under conflict for social event participation
Xiangyu Zhao, Tong Xu, Qi Liu, and Hao Guo. 2016 · 2016
Earlier work this paper cites.
Modeling temporal-spatial correlations for crime prediction
Xiangyu Zhao and Jiliang Tang. 2017 · 2017
Earlier work this paper cites.
Incorporating spatio-temporal smoothness for air quality inference
Xiangyu Zhao, Tong Xu, Yanjie Fu, Enhong Chen, and Hao Guo. 2017 · 2017
Earlier work this paper cites.
Geolocation Estimation of Photos using a Hierarchical Model and Scene Classification
Eric Müller-Budack, Kader Pustu-Iren, and Ralph Ewerth. 2018 · 2018
Earlier work this paper cites.
CPlaNet: Enhancing Image Geolocalization by Combinatorial Partitioning of Maps
Paul Hongsuck Seo, Tobias Weyand, Jack Sim, and Bohyung Han. 2018 · 2018
Earlier work this paper cites.
Lending orientation to neural networks for cross-view geo-localization
Liu Liu and Hongdong Li. 2019 · 2019
Earlier work this paper cites.
A Survey on Map-Based Localization Techniques for Autonomous Vehicles
Athanasios Chalvatzaras, Ioannis Pratikakis, and Angelos A Amanatiadis. 2022 · 2022
Earlier work this paper cites.
OPTDP: Towards optimal personalized trajectory differential privacy for trajectory data publishing
Wenqing Cheng, Ruxue Wen, Haojun Huang, Wang Miao, and Chen Wang. 2022 · 2022
Earlier work this paper cites.
G3: Geolocation via guidebook grounding
Grace Luo, Giscard Biamby, Trevor Darrell, Daniel Fried, and Anna Rohrbach. 2022 · 2022
Earlier work this paper cites.
Where in the World is this Image? Transformer-based Geo-localization in the Wild
Shraman Pramanick, Ewa M Nowara, Joshua Gleason, Carlos D Castillo, and Rama Chellappa. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022 · 2022
Cited alongside, same era.
Multi-type urban crime prediction
Xiangyu Zhao, Wenqi Fan, Hui Liu, and Jiliang Tang. 2022 · 2022
Cited alongside, same era.
TransGeo: Transformer Is All You Need for Cross-view Image Geo-localization
Sijie Zhu, Mubarak Shah, and Chen Chen. 2022 · 2022
Cited alongside, same era.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, and 1 others. 2023 · 2023
Cited alongside, same era.
Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, and 1 others. 2024 · 2024
Later among the works it cites.
Debunking war information disorder: A case study in assessing the use of multimedia verification tools
Sohail Ahmed Khan, Laurence Dierickx, Jan-Gunnar Furuly, Henrik Brattli Vold, Rano Tahseen, Carl-Gustav Linden, and Duc-Tien Dang-Nguyen. 2024 · 2024
Later among the works it cites.
Georeasoner: Geo-localization with reasoning in street views using a large vision-language model
Ling Li, Yu Ye, Bingchuan Jiang, and Wei Zeng. 2024 · 2024
Later among the works it cites.
Deepseek-vl: towards real-world vision-language understanding
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Yaofeng Sun, and 1 others. 2024 · 2024
Later among the works it cites.
Mm1: Methods, analysis & insights from multimodal llm pre-training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rizhao Cai, Zirui Song, Dayan Guan, Zhenhao Chen, Xing Luo, Chenyu Yi, and Alex Kot. 2023 · 2023
Cited alongside, same era.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023 · 2023
Cited alongside, same era.
Learning generalized zero-shot learners for open-domain image geolocalization
Lukas Haas, Silas Alberti, and Michal Skreta. 2023 · 2023
Cited alongside, same era.
Mitigating action hysteresis in traffic signal control with traffic predictive reinforcement learning
Xiao Han, Xiangyu Zhao, Liang Zhang, and Wanyu Wang. 2023 · 2023
Cited alongside, same era.
Gpt-4v dataset
LAION. 2023 · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. 2023 · 2023
Cited alongside, same era.
Aligning large multimodal models with factually augmented rlhf
Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, and 1 others. 2023 · 2023
Cited alongside, same era.
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Floris Weers, and 1 others. 2024 · 2024
Later among the works it cites.
Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
AI Meta. 2024 · 2024
Later among the works it cites.
OpenAI. 2024 · 2024
Later among the works it cites.
Geoclip: Clip-inspired alignment between locations and images for effective worldwide geo-localization
Vicente Vivanco Cepeda, Gaurav Kumar Nayak, and Mubarak Shah. 2024 · 2024
Later among the works it cites.
Universal adversarial perturbations for vision-language pre-trained models
Peng-Fei Zhang, Zi Huang, and Guangdong Bai. 2024 · 2024
Later among the works it cites.
Cobra: Extending mamba to multi-modal large language model for efficient inference
Han Zhao, Min Zhang, Wei Zhao, Pengxiang Ding, Siteng Huang, and Donglin Wang. 2024 · 2024
Later among the works it cites.
Tinyllava
Baichuan Zhou and Lei Huang. 2024 · 2024
Later among the works it cites.
Img2loc: Revisiting image geolocalization using multi-modality foundation models and image-based retrieval-augmented generation
Zhongliang Zhou, Jielu Zhang, Zihan Guan, Mengxuan Hu, Ni Lao, Lan Mu, Sheng Li, and Gengchen Mai. 2024 · 2024
Later among the works it cites.
Recognition through reasoning: Reinforcing image geo-localization with large vision-language models
Ling Li, Yao Zhou, Yuxuan Liang, Fugee Tsung, and Jiaheng Wei. 2025 · 2025
Closest in time.
Kimi k2: Open agentic intelligence
Kimi Team, Yifan Bai, Yiping Bao, Guanduo Chen, Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, and 1 others. 2025 · 2025
Closest in time.