Fetching the paper…
Reading the bibliography…
Cross-view geo-localization identifies the locations of street-view images by matching them with geo-tagged satellite images or OSM.
Historical review of ocr research and development
Shunji Mori, Ching Y Suen, and Kazuhiko Yamamoto · 1992
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
From structure-from-motion point clouds to fast location recognition
Arnold Irschara, Christopher Zach, Jan-Michael Frahm, and Horst Bischof · 2009
Earlier work this paper cites.
Building rome on a cloudless day
Jan-Michael Frahm, Pierre Fite-Georgel, David Gallup, Tim Johnson, Rahul Raguram, Changchang Wu, Yi-Hung Jen, Enrique Dunn, Brian Clipp, Svetlana Lazebnik, et al · 2010
Earlier work this paper cites.
Building rome in a day
Sameer Agarwal, Yasutaka Furukawa, Noah Snavely, Ian Simon, Brian Curless, Steven M Seitz, and Richard Szeliski · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Wide-area image geolocalization with aerial reference imagery
Scott Workman, Richard Souvenir, and Nathan Jacobs · 2015
Earlier work this paper cites.
Openmvg: Open multiple view geometry
Pierre Moulon, Pascal Monasse, Romuald Perrot, and Renaud Marlet · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel · 2018
Earlier work this paper cites.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi · 2019
Earlier work this paper cites.
Visual semantic reasoning for image-text matching
Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu · 2019
Earlier work this paper cites.
Lending orientation to neural networks for cross-view geo-localization
Liu Liu and Hongdong Li · 2019
Earlier work this paper cites.
From coarse to fine: Robust hierarchical localization at large scale
Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk · 2019
Earlier work this paper cites.
Spatial-aware feature aggregation for image based cross-view geo-localization
Yujiao Shi, Liu Liu, Xin Yu, and Hongdong Li · 2019
Earlier work this paper cites.
Large-scale, real-time visual–inertial localization revisited
Simon Lynen, Bernhard Zeisl, Dror Aiger, Michael Bosse, Joel Hesch, Marc Pollefeys, Roland Siegwart, and Torsten Sattler · 2020
Earlier work this paper cites.
Where am i looking at? joint location and orientation estimation by cross-view matching
Yujiao Shi, Xin Yu, Dylan Campbell, and Hongdong Li · 2020
Earlier work this paper cites.
Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers
Hila Chefer, Shir Gur, and Lior Wolf · 2021
Cited alongside, same era.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Vpc-net: Completion of 3d vehicles from mls point clouds
Yan Xia, Yusheng Xu, Cheng Wang, and Uwe Stilla · 2021
Cited alongside, same era.
Vigor: Cross-view image geo-localization beyond one-to-one retrieval
Sijie Zhu, Taojiannan Yang, and Chen Chen · 2021
Cited alongside, same era.
An empirical study of training end-to-end vision-and-language transformers
A survey of hallucination in large foundation models
Vipula Rawte, Amit Sheth, and Amitava Das · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
Towards unified text-based person retrieval: A large-scale multi-attribute and language search benchmark
Shuyu Yang, Yinan Zhou, Zhedong Zheng, Yaxiong Wang, Li Zhu, and Yujiao Wu · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Later among the works it cites.
Cross-view geo-localization via learning disentangled geometric layout correspondence
Xiaohan Zhang, Xingyu Li, Waqas Sultani, Yi Zhou, and Safwan Wshah · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Pengchuan Zhang, Lu Yuan, Nanyun Peng, et al · 2022
Cited alongside, same era.
Text2pos: Text-to-point-cloud cross-modal localization
Manuel Kolmet, Qunjie Zhou, Aljoša Ošep, and Laura Leal-Taixé · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Cited alongside, same era.
Open world entity segmentation
Lu Qi, Jason Kuen, Yi Wang, Jiuxiang Gu, Hengshuang Zhao, Philip Torr, Zhe Lin, and Jiaya Jia · 2022
Cited alongside, same era.
Multi-grained vision language pre-training: Aligning texts with visual concepts
Yan Zeng, Xinsong Zhang, and Hang Li · 2022
Cited alongside, same era.
Sharegpt4v: Improving large multi-modal models with better captions
Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin · 2023
Cited alongside, same era.
Towards natural language-guided drones: Geotext-1652 benchmark with spatially relation matching
Meng Chu, Zhedong Zheng, Wei Ji, and Tat-Seng Chua · 2023
Cited alongside, same era.
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic · 2024
Closest in time.
Latteclip: Unsupervised clip fine-tuning via lmm-synthetic texts
Anh-Quan Cao, Maximilian Jaritz, Matthieu Guillaumin, Raoul de Charette, and Loris Bazzani · 2024
Closest in time.
” where am i?” scene retrieval with language
Jiaqi Chen, Daniel Barath, Iro Armeni, Marc Pollefeys, and Hermann Blum · 2024
Closest in time.
Eva-02: A visual representation for neon genesis
Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao · 2024
Closest in time.
Hello gpt-4o
OpenAI · 2024
Closest in time.
Loc4plan: Locating before planning for outdoor vision and language navigation
Huilin Tian, Jingke Meng, Wei-Shi Zheng, Yuan-Ming Li, Junkai Yan, and Yunong Zhang · 2024
Closest in time.
Fine-grained cross-view geo-localization using a correlation-aware homography estimator
Xiaolong Wang, Runsen Xu, Zhuofan Cui, Zeyu Wan, and Yu Zhang · 2024
Closest in time.
Text2loc: 3d point cloud localization from natural language
Yan Xia, Letian Shi, Zifeng Ding, Joao F Henriques, and Daniel Cremers · 2024
Closest in time.
Adapting fine-grained cross-view localization to areas without fine ground truth
Zimin Xia, Yujiao Shi, Hongdong Li, and Julian FP Kooij · 2025
Closest in time.
Long-clip: Unlocking the long-text capability of clip
Beichen Zhang, Pan Zhang, Xiaoyi Dong, Yuhang Zang, and Jiaqi Wang · 2025
Closest in time.
Open panoramic segmentation
Junwei Zheng, Ruiping Liu, Yufan Chen, Kunyu Peng, Chengzhi Wu, Kailun Yang, Jiaming Zhang, and Rainer Stiefelhagen · 2025
Closest in time.