Fetching the paper…
Reading the bibliography…
Recent evaluations of Large Multimodal Models (LMMs) have explored their capabilities in various domains, with only few benchmarks specifically focusing on urban environments.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
IM2GPS: estimating geographic information from a single image
Hays, J.; and Efros, A. A. 2008 · 2008
Earlier work this paper cites.
The Cityscapes Dataset for Semantic Urban Scene Understanding
Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; and Schiele, B. 2016 · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016 · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
DOTA: A large-scale dataset for object detection in aerial images
Xia, G.-S.; Bai, X.; Ding, J.; Zhu, Z.; Belongie, S.; Luo, J.; Datcu, M.; Pelillo, M.; and Zhang, L. 2018 · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A.; and Manning, C. D. 2019 · 2019
Earlier work this paper cites.
The mapillary traffic sign dataset for detection and classification on a global scale
Ertler, C.; Mislej, J.; Ollmann, T.; Porzi, L.; Neuhold, G.; and Kuang, Y. 2020 · 2020
Earlier work this paper cites.
RSVQA: Visual question answering for remote sensing data
Lobry, S.; Marcos, D.; Murray, J.; and Tuia, D. 2020 · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Learning to count everything
Ranjan, V.; Sharma, U.; Nguyen, T.; and Hoai, M. 2021 · 2021
Earlier work this paper cites.
Vigor: Cross-view image geo-localization beyond one-to-one retrieval
Zhu, S.; Yang, T.; and Chen, C. 2021 · 2021
Earlier work this paper cites.
Beyond cross-view image retrieval: Highly accurate vehicle localization using satellite image
Shi, Y.; and Li, H. 2022 · 2022
Earlier work this paper cites.
FAIR1M: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery
Sun, X.; Wang, P.; Yan, Z.; Xu, F.; Wang, R.; Diao, W.; Chen, J.; Li, J.; Feng, Y.; Xu, T.; et al. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Fu, C.; Chen, P.; Shen, Y.; Qin, Y.; Zhang, M.; Lin, X.; Yang, J.; Zheng, X.; Li, K.; Sun, X.; et al. 2023 · 2023
Cited alongside, same era.
Rsgpt: A remote sensing vision language model and benchmark
Hu, Y.; Yuan, J.; Wen, C.; Lu, X.; and Li, X. 2023 · 2023
Cited alongside, same era.
Geochat: Grounded large vision-language model for remote sensing
Kuckreja, K.; Danish, M. S.; Naseer, M.; Das, A.; Khan, S.; and Khan, F. S. 2024 · 2024
Closest in time.
What matters when building vision-language models?
Laurençon, H.; Tronchon, L.; Cord, M.; and Sanh, V. 2024 · 2024
Closest in time.
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Li, F.; Zhang, R.; Zhang, H.; Zhang, Y.; Li, B.; Li, W.; Ma, Z.; and Li, C. 2024 · 2024
Closest in time.
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
Li, X.; Ding, J.; and Elhoseiny, M. 2024 · 2024
Closest in time.
Vila: On pre-training for visual language models
Lin, J.; Yin, H.; Ping, W.; Molchanov, P.; Shoeybi, M.; and Han, S. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Cited alongside, same era.
A survey on multimodal large language models
Yin, S.; Fu, C.; Zhao, S.; Li, K.; Sun, X.; Xu, T.; and Chen, E. 2023 · 2023
Cited alongside, same era.
Zhang, P.; Wang, X. D. B.; Cao, Y.; Xu, C.; Ouyang, L.; Zhao, Z.; Ding, S.; Zhang, S.; Duan, H.; Yan, H.; et al. 2023 · 2023
Cited alongside, same era.
The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic. 2024 · 2024
Cited alongside, same era.
A Survey of Multimodal Large Language Model from A Data-centric Perspective
Bai, T.; Liang, H.; Wan, B.; Yang, L.; Li, B.; Wang, Y.; Cui, B.; He, C.; Yuan, B.; and Zhang, W. 2024 · 2024
Cited alongside, same era.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Chen, Z.; Wang, W.; Tian, H.; Ye, S.; Gao, Z.; Cui, E.; Tong, W.; Hu, K.; Luo, J.; Ma, Z.; et al. 2024 · 2024
Cited alongside, same era.
A survey on multimodal large language models for autonomous driving
Cui, C.; Ma, Y.; Cao, X.; Ye, W.; Zhou, Y.; Liang, K.; Chen, J.; Lu, J.; Yang, Z.; Liao, K.-D.; et al. 2024 · 2024
Cited alongside, same era.
Hao, X.; Chen, W.; Yan, Y.; Zhong, S.; Wang, K.; Wen, Q.; and Liang, Y. 2024 · 2024
Cited alongside, same era.
Closest in time.
Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model
Muhtar, D.; Li, Z.; Gu, F.; Zhang, X.; and Xiao, P. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M.; Savinov, N.; Teplyashin, D.; Lepikhin, D.; Lillicrap, T.; Alayrac, J.-b.; Soricut, R.; Lazaridou, A.; Firat, O.; Schrittwieser, J.; et al. 2024 · 2024
Closest in time.
Velma: Verbalization embodiment of llm agents for vision and language navigation in street view
Schumann, R.; Zhu, W.; Feng, W.; Fu, T.-J.; Riezler, S.; and Wang, W. Y. 2024 · 2024
Closest in time.
Urban Population (% of Total Population) - World
World Bank. 2024 · 2024
Closest in time.
A comprehensive survey of large language models and multimodal large language models in medicine
Xiao, H.; Zhou, F.; Liu, X.; Liu, T.; Li, Z.; Liu, X.; and Huang, X. 2024 · 2024
Closest in time.
Urbanclip: Learning text-enhanced urban region profiling with contrastive language-image pretraining from the web
Yan, Y.; Wen, H.; Zhong, S.; Chen, W.; Chen, H.; Wen, Q.; Zimmermann, R.; and Liang, Y. 2024 · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X.; Ni, Y.; Zhang, K.; Zheng, T.; Liu, R.; Zhang, G.; Stevens, S.; Jiang, D.; Ren, W.; Sun, Y.; et al. 2024 · 2024
Closest in time.