Fetching the paper…
Reading the bibliography…
Previous methods for image geo-localization have typically treated the task as either classification or retrieval, often relying on black-box decisions that lack interpretability.
Urban planning and gis
Anthony GO Yeh · 1999
Earlier work this paper cites.
Vision for mobile robot navigation: A survey
Guilherme N DeSouza and Avinash C Kak · 2002
Earlier work this paper cites.
Im2gps: estimating geographic information from a single image
James Hays and Alexei A Efros · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Cybercasing the joint: On the privacy implications of { \{ Geo-Tagging } \}
Gerald Friedland, Robin Sommer, et al · 2010
Earlier work this paper cites.
Geodma—geographic data mining analyst
Thales Sehn Körting, Leila Maria Garcia Fonseca, and Gilberto Câmara · 2013
Earlier work this paper cites.
Planet-photo geolocation with convolutional neural networks
Tobias Weyand, Ilya Kostrikov, and James Philbin · 2016
Earlier work this paper cites.
Netvlad: Cnn architecture for weakly supervised place recognition
Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic · 2016
Earlier work this paper cites.
The benchmarking initiative for multimedia evaluation: Mediaeval 2016
Martha Larson, Mohammad Soleymani, Guillaume Gravier, Bogdan Ionescu, and Gareth JF Jones · 2017
Earlier work this paper cites.
Webvision database: Visual learning and understanding from web data
Wen Li, Limin Wang, Wei Li, Eirikur Agustsson, and Luc Van Gool · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Revisiting im2gps in the deep learning era
Nam Vo, Nathan Jacobs, and James Hays · 2017
Earlier work this paper cites.
Cplanet: Enhancing image geolocalization by combinatorial partitioning of maps
Paul Hongsuck Seo, Tobias Weyand, Jack Sim, and Bohyung Han · 2018
Earlier work this paper cites.
Geolocation estimation of photos using a hierarchical model and scene classification
Eric Müller-Budack, Kader Pustu-Iren, and Ralph Ewerth · 2018
Earlier work this paper cites.
Geoman: Multi-level attention networks for geo-sensory time series prediction
Yuxuan Liang, Songyu Ke, Junbo Zhang, Xiuwen Yi, and Yu Zheng · 2018
Earlier work this paper cites.
Measuring daily accessed street greenery: A human-scale approach for informing better urban planning practices
Yu Ye, Daniel Richards, Yi Lu, Xiaoping Song, Yu Zhuang, Wei Zeng, and Teng Zhong · 2019
Earlier work this paper cites.
The visual quality of streets: A human-centred continuous measurement based on machine learning algorithms and street view images
Yu Ye, Wei Zeng, Qiaomu Shen, Xiaohu Zhang, and Yi Lu · 2019
Earlier work this paper cites.
Urban traffic prediction from spatio-temporal data using deep meta learning
Zheyi Pan, Yuxuan Liang, Weifeng Wang, Yong Yu, Yu Zheng, and Junbo Zhang · 2019
Earlier work this paper cites.
Exploiting the earth’s spherical geometry to geolocate images
Mike Izbicki, Evangelos E Papalexakis, and Vassilis J Tsotras · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Cross-view geo-localization with layer-to-layer transformer
Hongji Yang, Xiufan Lu, and Yingying Zhu · 2021
Earlier work this paper cites.
Per-pixel classification is not all you need for semantic segmentation
Bowen Cheng, Alex Schwing, and Alexander Kirillov · 2021
Earlier work this paper cites.
Where in the world is this image? transformer-based geo-localization in the wild
Shraman Pramanick, Ewa M Nowara, Joshua Gleason, Carlos D Castillo, and Rama Chellappa · 2022
Earlier work this paper cites.
Transgeo: Transformer is all you need for cross-view image geo-localization
Sijie Zhu, Mubarak Shah, and Chen Chen · 2022
Earlier work this paper cites.
Rethinking visual geo-localization for large-scale applications
Gabriele Berton, Carlo Masone, and Barbara Caputo · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Earlier work this paper cites.
Learning with noisy labels revisited: A study using real-world human annotations
Jiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu, Gang Niu, and Yang Liu · 2022
Earlier work this paper cites.
Core challenges of social robot navigation: A survey
Christoforos Mavrogiannis, Francesca Baldini, Allan Wang, Dapeng Zhao, Pete Trautman, Aaron Steinfeld, and Jean Oh · 2023
Earlier work this paper cites.
Where we are and what we’re looking at: Query-based worldwide image geo-localization using hierarchies and scenes
Brandon Clark, Alec Kerrigan, Parth Parag Kulkarni, Vicente Vivanco Cepeda, and Mubarak Shah · 2023
Earlier work this paper cites.
Fine-grained cross-view geo-localization using a correlation-aware homography estimator
Xiaolong Wang, Runsen Xu, Zhuofan Cui, Zeyu Wan, and Yu Zhang · 2023
Cited alongside, same era.
Geoclip: Clip-inspired alignment between locations and images for effective worldwide geo-localization
Vicente Vivanco Cepeda, Gaurav Kumar Nayak, and Mubarak Shah · 2023
Cited alongside, same era.
Learning generalized zero-shot learners for open-domain image geolocalization
Lukas Haas, Silas Alberti, and Michal Skreta · 2023
Cited alongside, same era.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2024
Later among the works it cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao · 2024
Later among the works it cites.
Seed-bench: Benchmarking multimodal large language models
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan · 2024
Later among the works it cites.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al · 2024
Later among the works it cites.
Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold
Amrith Setlur, Saurabh Garg, Xinyang Geng, Naman Garg, Virginia Smith, and Aviral Kumar · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Cited alongside, same era.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu · 2023
Cited alongside, same era.
Mathprompter: Mathematical reasoning using large language models
Shima Imani, Liang Du, and Harsh Shrivastava · 2023
Cited alongside, same era.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Cited alongside, same era.
Later among the works it cites.
Hallucination of multimodal large language models: A survey
Zechen Bai, Pichao Wang, Tianjun Xiao, Tong He, Zongbo Han, Zheng Zhang, and Mike Zheng Shou · 2024
Later among the works it cites.
Measuring and reducing llm hallucination without gold-standard answers
Jiaheng Wei, Yuanshun Yao, Jean-Francois Ton, Hongyi Guo, Andrew Estornell, and Yang Liu · 2024
Later among the works it cites.
Fvel: Interactive formal verification environment with large language models via theorem proving
Xiaohan Lin, Qingxing Cao, Yinya Huang, Haiming Wang, Jianqiao Lu, Zhengying Liu, Linqi Song, and Xiaodan Liang · 2024
Later among the works it cites.
Llm-check: Investigating detection of hallucinations in large language models
Gaurang Sriramanan, Siddhant Bharti, Vinu Sankar Sadasivan, Shoumik Saha, Priyatham Kattakinda, and Soheil Feizi · 2024
Later among the works it cites.
Erbench: An entity-relationship based automatically verifiable hallucination benchmark for large language models
Jio Oh, Soyeon Kim, Junseok Seo, Jindong Wang, Ruochen Xu, Xing Xie, and Steven Whang · 2024
Later among the works it cites.
Automatic dataset construction (adc): Sample collection, data curation, and beyond
Minghao Liu, Zonglin Di, Jiaheng Wei, Zhongruo Wang, Hengxiang Zhang, Ruixuan Xiao, Haoyu Wang, Jinlong Pang, Hao Chen, Ankit Shah, et al · 2024
Later among the works it cites.
Mtms: Multi-teacher multi-stage knowledge distillation for reasoning-based machine reading comprehension
Zhuo Zhao, Zhiwen Xie, Guangyou Zhou, and Jimmy Xiangji Huang · 2024
Later among the works it cites.
Openstreetview-5m: The many roads to global visual geolocation
Guillaume Astruc, Nicolas Dufour, Ioannis Siglidis, Constantin Aronssohn, Nacim Bouia, Stephanie Fu, Romain Loiseau, Van Nguyen Nguyen, Charles Raude, Elliot Vincent, et al · 2024
Later among the works it cites.
A comprehensive review on autonomous navigation
Saeid Nahavandi, Roohallah Alizadehsani, Darius Nahavandi, Shady Mohamed, Navid Mohajer, Mohammad Rokonuzzaman, and Ibrahim Hossain · 2025
Closest in time.
F G 2 {FG}^{2} : Fine-grained cross-view localization by fine-grained feature matching
Zimin Xia and Alexandre Alahi · 2025
Closest in time.
Georanker: Distance-aware ranking for worldwide image geolocalization
Pengyue Jia, Seongheon Park, Song Gao, Xiangyu Zhao, and Yixuan Li · 2025
Closest in time.
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin · 2025
Closest in time.
Geolocation with real human gameplay data: A large-scale dataset and human-like reasoning framework
Zirui Song, Jingpu Yang, Yuan Huang, Jonathan Tonglet, Zeyu Zhang, Tao Cheng, Meng Fang, Iryna Gurevych, and Xiuying Chen · 2025
Closest in time.
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models
Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Yuchen Duan, Hao Tian, Weijie Su, Jie Shao, et al · 2025
Closest in time.
Nature makes no leaps: Building continuous location embeddings with satellite imagery from the web
Xixuan Hao, Wei Chen, Xingchen Zou, and Yuxuan Liang · 2025
Closest in time.
Around the world in 80 timesteps: A generative approach to global visual geolocation
Nicolas Dufour, Vicky Kalogeiton, David Picard, and Loic Landrieu · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Skywork open reasoner 1 technical report
Jujie He, Jiacai Liu, Chris Yuhao Liu, Rui Yan, Chaojie Wang, Peng Cheng, Xiaoyu Zhang, Fuxiang Zhang, Jiacheng Xu, Wei Shen, et al · 2025
Closest in time.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al · 2025
Closest in time.
Reinforced mllm: A survey on rl-based reasoning in multimodal large language models
Guanghao Zhou, Panjia Qiu, Cen Chen, Jie Wang, Zheming Yang, Jian Xu, and Minghui Qiu · 2025
Closest in time.
Gaea: A geolocation aware conversational model
Ron Campos, Ashmal Vayani, Parth Parag Kulkarni, Rohit Gupta, Aritra Dutta, and Mubarak Shah · 2025
Closest in time.
Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, et al · 2025
Closest in time.
Dong Guo, Faming Wu, Feida Zhu, Fuxing Leng, Guang Shi, Haobin Chen, Haoqi Fan, Jian Wang, Jianyu Jiang, Jiawei Wang, et al · 2025
Closest in time.