Fetching the paper…
Reading the bibliography…
Image geolocation is a critical task in various image-understanding applications.
Im2gps: estimating geographic information from a single image
James Hays and Alexei A Efros · 2008
Earlier work this paper cites.
Deepgeo: Photo localization with deep neural network
Sudharshan Suresh, Nathaniel Chodosh, and Montiel Abello · 2018
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Im2city: image geo-localization via multi-modal learning
Meiliu Wu and Qunying Huang · 2022
Earlier work this paper cites.
Visual commonsense in pretrained unimodal and multimodal models
Chenyu Zhang, Benjamin Van Durme, Zhuowan Li, and Elias Stengel-Eskin · 2022
Earlier work this paper cites.
Vicente Vivanco Cepeda, Gaurav Kumar Nayak, and Mubarak Shah · 2023
Cited alongside, same era.
Pigeon: Predicting image geolocations
Lukas Haas, Silas Alberti, and Michal Skreta · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao · 2023
Cited alongside, same era.
Large language models are zero-shot text classifiers
Zhiqiang Wang, Yiran Pang, and Yanbin Lin · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2023
Later among the works it cites.
Fuyu-8b: A multimodal architecture for ai agents, 2024
Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, and Sağnak Taşırlar · 2024
Closest in time.
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Xilin Wei, Songyang Zhang, Haodong Duan, Maosong Cao, et al · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Cited alongside, same era.
Improved baselines with visual instruction tuning, 2023a
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee
Cited in the paper.
Visual instruction tuning, 2023b
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee
Cited in the paper.
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee · 2024
Closest in time.
GPT-4V(ision) System Card, 2023
OpenAI · 2024
Closest in time.