Fetching the paper…
Reading the bibliography…
This paper introduces GeoChain, a large-scale benchmark for evaluating step-by-step geographic reasoning in multimodal large language models (MLLMs).
Cross-view image matching for geo-localization in urban environments
Yicong Tian, Chen Chen, and Mubarak Shah. 2017 · 2006
Earlier work this paper cites.
Openstreetmap: User-generated street maps
Muki Haklay and Patrick Weber. 2008 · 2008
Earlier work this paper cites.
IM2GPS: estimating geographic information from a single image
James Hays and Alexei A. Efros. 2008 · 2008
Earlier work this paper cites.
Google street view: Capturing the world at street level
Dragomir Anguelov, Carole Dulong, Daniel Filip, Christian Frueh, Stéphane Lafon, Richard Lyon, Abhijit Ogale, Luc Vincent, and Josh Weaver. 2010 · 2010
Earlier work this paper cites.
Geoguessr - let’s explore the world!
GeoGuessr. 2013 · 2013
Earlier work this paper cites.
NetVLAD: CNN architecture for weakly supervised place recognition
Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic. 2016 · 2016
Earlier work this paper cites.
PlaNet - photo geolocation with convolutional neural networks
Tobias Weyand, Ilya Kostrikov, and James Philbin. 2016 · 2016
Earlier work this paper cites.
Scene parsing through ADE20K dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. 2017 · 2017
Earlier work this paper cites.
Object hallucination in image captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018 · 2018
Earlier work this paper cites.
Generative deep-neural-network mixture modeling with semi-supervised minmax+em learning
Nilay Pande and Suyash P. Awate. 2021 · 2020
Earlier work this paper cites.
Mapillary street-level sequences: A dataset for lifelong place recognition
Frederik Warburg, Søren Hauberg, Gregory D. D. Funke, and Yoko Yuki. 2020 · 2020
Cited alongside, same era.
Per-pixel classification is not all you need for semantic segmentation
Bowen Cheng, Alexander G. Schwing, and Alexander Kirillov. 2021 · 2021
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023 · 2023
Cited alongside, same era.
Maea: Multimodal attribution for embodied ai
Vidhi Jain, Jayant Sravan Tamarapalli, Sahiti Yerramilli, and Yonatan Bisk. 2023 · 2023
Cited alongside, same era.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024 · 2024
Later among the works it cites.
VLMs as GeoGuessr masters—exceptional performance, hidden biases, and privacy risks
Ruixiang Yang, Cheng Zhang, Lingxi Meng, He Wang, Xiaoyan Li, Yuke Li, Shuo Wang, Haoran Wei, Yiyang Li, Wentao Qu, Pengchuan Zhang, Jiazheng Xu, Bihan Wen, Diyi Yang, Kangkang Lu, Saurabh Gupta, Guanzhong Wang, Zhiqiang Shen, Baining Guo, and 3 others. 2024 · 2024
Later among the works it cites.
GAEA: A geolocation aware conversational model
Ron Campos, Ashmal Vayani, Parth Parag Kulkarni, Rohit Gupta, Aritra Dutta, and Mubarak Shah. 2025 · 2025
Closest in time.
Gemini 2.5 flash is now in preview
Google. 2025a · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexandre Lacoste, Nils Lehmann, Pau Rodriguez, Evan David Sherwin, Hannah Kerner, Björn Lütjens, Jeremy Andrew Irvin, David Dao, Hamed Alemohammad, Alexandre Drouin, Mehmet Gunturkun, Gabriel Huang, David Vazquez, Dava Newman, Yoshua Bengio, Stefano Ermon, and Xiao Xiang Zhu. 2023 · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen. 2023 · 2023
Cited alongside, same era.
Georeasoner: Geo-localization with reasoning in street views using a large vision-language model
Ling Li, Yu Ye, Bingchuan Jiang, and Wei Zeng. 2024 · 2024
Cited alongside, same era.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, and 262 others. 2024 · 2024
Cited alongside, same era.
Evaluating precise geolocation inference capabilities of vision language models
Saksham Pramanik, Aayush Mundra, Ashutosh Mittal, Sreyas Mohan, Saim Wani, Shramay S Vernekar, Pranav M Dixit, Shanti Priya, Ankur Beniwal, Ojaswa Sharma, and Senthil Mani. 2024 · 2024
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, Soroosh Mariooryad, Yifan Ding, Xinyang Geng, Fred Alcober, Roy Frostig, Mark Omernick, Lexi Walker, Cosmin Paduraru, Christina Sorokin, and 1118 others. 2024 · 2024
Cited alongside, same era.
Embodied symbiotic assistants that see, act, infer and chat
Yuchen Cao, Nilay Pande, Ayush Jain, Shikhar Sharma, Gabriel Sarch, Nikolaos Gkanatsios, Xian Zhou, and Katerina Fragkiadaki
Cited in the paper.
Attribution regularization for multimodal paradigms
Sahiti Yerramilli, Jayant Sravan Tamarapalli, Jonathan Francis, and Eric Nyberg. 2024a
Cited in the paper.
Google. 2025b · 2025
Closest in time.
Huemanity: Probing fine-grained visual perception in mllms
Rynaa Grover, Jayant Sravan Tamarapalli, Sahiti Yerramilli, and Nilay Pande. 2025 · 2025
Closest in time.
Ai guide dog: Egocentric path prediction on smartphone
Aishwarya Jadhav, Jeffery Cao, Abhishree Shetty, Urvashi Kumar, Aditi Sharma, Ben Sukboontip, Jayant Tamarapalli, Jingyi Zhang, and Aniruddh Koul. 2025 · 2025
Closest in time.
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, Adam Khoja, Ryan Kim, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, and 1090 others. 2025 · 2025
Closest in time.
Geolocation with real human gameplay data: A large-scale dataset and human-like reasoning framework
Zirui Song, Jingpu Yang, Yuan Huang, Jonathan Tonglet, Zeyu Zhang, Tao Cheng, Meng Fang, Iryna Gurevych, and Xiuying Chen. 2025 · 2025
Closest in time.
Claude 3.7 sonnet documentation
Claude 3.7 Sonnet. 2025 · 2025
Closest in time.