Fetching the paper…
Reading the bibliography…
Visual-Language Models (VLMs) have shown remarkable performance across various tasks, particularly in recognizing geographic information from images.
A mathematical theory of communication
Claude E Shannon. 1948 · 1948
Earlier work this paper cites.
Attitudinal effects of mere exposure
Robert B Zajonc. 1968 · 1968
Earlier work this paper cites.
Using deep learning and google street view to estimate the demographic makeup of neighborhoods across the united states
Timnit Gebru, Jonathan Krause, Yilun Wang, Duyun Chen, Jia Deng, Erez Lieberman Aiden, and Li Fei-Fei. 2017 · 2017
Earlier work this paper cites.
The echo chamber effect on social media
Matteo Cinelli, Gianmarco De Francisci Morales, Alessandro Galeazzi, Walter Quattrociocchi, and Michele Starnini. 2021 · 2021
Earlier work this paper cites.
Analyzing the effects of green view index of neighborhood streets on walking time using google street view and deep learning
Donghwan Ki and Sugie Lee. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, and 1 others. 2021 · 2021
Earlier work this paper cites.
Measuring social biases in grounded vision and language embeddings
Candace Ross, Boris Katz, and Andrei Barbu. 2021 · 2021
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Earlier work this paper cites.
Worst of both worlds: Biases compound in pre-trained vision-and-language models
Tejas Srinivasan and Yonatan Bisk. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022 · 2022
Earlier work this paper cites.
American== white in multimodal language-and-image ai
Robert Wolfe and Aylin Caliskan. 2022 · 2022
Earlier work this paper cites.
Counterfactually measuring and eliminating social bias in vision-language pre-training models
Yi Zhang, Junyang Wang, and Jitao Sang. 2022 · 2022
Earlier work this paper cites.
Foundational models in medical imaging: A comprehensive survey and future vision
Bobby Azad, Reza Azad, Sania Eskandari, Afshin Bozorgpour, Amirhossein Kazerouni, Islem Rekik, and Dorit Merhof. 2023 · 2023
Earlier work this paper cites.
Are large language models geospatially knowledgeable?
Prabin Bhandari, Antonios Anastasopoulos, and Dieter Pfoser. 2023 · 2023
Earlier work this paper cites.
A multidimensional analysis of social biases in vision transformers
Jannik Brinkmann, Paul Swoboda, and Christian Bartelt. 2023 · 2023
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, and 1 others. 2023 · 2023
Earlier work this paper cites.
Multimodal foundation models exploit text to make medical image predictions
Thomas Buckley, James A. Diao, Pranav Rajpurkar, Adam Rodman, and Arjun K. Manrai. 2023 · 2023
Earlier work this paper cites.
Urban visual intelligence: Uncovering hidden city profiles with street view images
Zhuangyuan Fan, Fan Zhang, Becky PY Loo, and Carlo Ratti. 2023 · 2023
Earlier work this paper cites.
‘person’== light-skinned, western man, and sexualization of women of color: Stereotypes in stable diffusion
Sourojit Ghosh and Aylin Caliskan. 2023 · 2023
Earlier work this paper cites.
Visogender: A dataset for benchmarking gender bias in image-text pronoun resolution
Siobhan Mackenzie Hall, Fernanda Gonçalves Abrantes, Hanwen Zhu, Grace Sodunke, Aleksandar Shtedritski, and Hannah Rose Kirk. 2023 · 2023
Cited alongside, same era.
Geo-knowledge-guided gpt models improve the extraction of location descriptions from disaster-related social media messages
Yingjie Hu, Gengchen Mai, Chris Cundy, Kristy Choi, Ni Lao, Wei Liu, Gaurish Lakhanpal, Ryan Zhenqi Zhou, and Kenneth Joseph. 2023 · 2023
Cited alongside, same era.
Geolm: Empowering language models for geospatially grounded language understanding
Zekun Li, Wenxuan Zhou, Yao-Yi Chiang, and Muhao Chen. 2023 · 2023
Cited alongside, same era.
Societal bias in vision-and-language datasets and models
Yuta Nakashima, Yusuke Hirota, Yankun Wu, and Noa Garcia. 2023 · 2023
Cited alongside, same era.
Gpt4geo: How a language model sees the world’s geography
Jonathan Roberts, Timo Lüddecke, Sowmen Das, Kai Han, and Samuel Albanie. 2023 · 2023
Biasdora: Exploring hidden biased associations in vision-language models
Chahat Raj, Anjishnu Mukherjee, Aylin Caliskan, Antonios Anastasopoulos, and Ziwei Zhu. 2024 · 2024
Later among the works it cites.
A unified framework and dataset for assessing societal bias in vision-language models
Ashutosh Sathe, Prachi Jain, and Sunayana Sitaram. 2024 · 2024
Later among the works it cites.
Assessment of multimodal large language models in alignment with human values
Zhelun Shi, Zhipin Wang, Hongxing Fan, Zaibin Zhang, Lijun Li, Yongting Zhang, Zhenfei Yin, Lu Sheng, Yu Qiao, and Jing Shao. 2024 · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, and 1 others. 2024 · 2024
Later among the works it cites.
New job, new gender? measuring the social bias in image generation models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A multi-dimensional study on bias in vision-language models
Gabriele Ruggeri and Debora Nozza. 2023 · 2023
Cited alongside, same era.
Biasasker: Measuring the bias in conversational ai system
Yuxuan Wan, Wenxuan Wang, Pinjia He, Jiazhen Gu, Haonan Bai, and Michael R Lyu. 2023 · 2023
Cited alongside, same era.
Contrastive language-vision ai models pretrained on web-scraped multimodal data exhibit sexual objectification bias
Robert Wolfe, Yiwei Yang, Bill Howe, and Aylin Caliskan. 2023 · 2023
Cited alongside, same era.
Geo-seq2seq: Twitter user geolocation on noisy data through sequence to sequence learning
Jingyu Zhang, Alexandra DeLucia, Chenyu Zhang, and Mark Dredze. 2023 · 2023
Cited alongside, same era.
K2: A foundation language model for geoscience knowledge understanding and utilization
Cheng Deng, Tianhang Zhang, Zhongmou He, Qiyuan Chen, Yuanyuan Shi, Yi Xu, Luoyi Fu, Weinan Zhang, Xinbing Wang, Chenghu Zhou, and 1 others. 2024 · 2024
Cited alongside, same era.
Examining gender and racial bias in large vision–language models using a novel dataset of parallel images
Kathleen C Fraser and Svetlana Kiritchenko. 2024 · 2024
Cited alongside, same era.
Pigeon: Predicting image geolocations
Lukas Haas, Michal Skreta, Silas Alberti, and Chelsea Finn. 2024 · 2024
Cited alongside, same era.
Wenxuan Wang, Haonan Bai, Jen-tse Huang, Yuxuan Wan, Youliang Yuan, Haoyi Qiu, Nanyun Peng, and Michael Lyu. 2024 · 2024
Later among the works it cites.
Comparing traditional and llm-based search for image geolocation
Albatool Wazzan, Stephen MacNeil, and Richard Souvenir. 2024 · 2024
Later among the works it cites.
Benchmarking trustworthiness of multimodal large language models: A comprehensive study
Yichi Zhang, Yao Huang, Yitong Sun, Chang Liu, Zhe Zhao, Zhengwei Fang, Yifan Wang, Huanran Chen, Xiao Yang, Xingxing Wei, and 1 others. 2024 · 2024
Later among the works it cites.
Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras
Abdelrahman Abouelenin, Atabak Ashfaq, Adam Atkinson, Hany Awadalla, Nguyen Bach, Jianmin Bao, Alon Benhaim, Martin Cai, Vishrav Chaudhary, Congcong Chen, and 1 others. 2025 · 2025
Closest in time.
Ocean-ocr: Towards general ocr application via a vision-language model
Song Chen, Xinyu Guo, Yadong Li, Tao Zhang, Mingan Lin, Dongdong Kuang, Youwei Zhang, Lingfeng Ming, Fengyu Zhang, Yuran Wang, and 1 others. 2025 · 2025
Closest in time.
Physbench: Benchmarking and enhancing vision-language models for physical world understanding
Wei Chow, Jiageng Mao, Boyi Li, Daniel Seita, Vitor Guizilini, and Yue Wang. 2025 · 2025
Closest in time.
Gender bias in large language models across multiple languages: A case study of chatgpt
Yitian Ding, Jinman Zhao, Chen Jia, Yining Wang, Zifan Qian, Weizhe Chen, and Xingyu Yue. 2025 · 2025
Closest in time.
Faircoder: Evaluating social bias of llms in code generation
Yongkang Du, Jen-tse Huang, Jieyu Zhao, and Lu Lin. 2025 · 2025
Closest in time.
Visbias: Measuring explicit and implicit social biases in vision language models
Jen-tse Huang, Jiantong Qin, Jianping Zhang, Youliang Yuan, Wenxuan Wang, and Jieyu Zhao. 2025a · 2025
Closest in time.
Where fact ends and fairness begins: Redefining ai bias evaluation through cognitive biases
Jen-tse Huang, Yuhang Yan, Linqi Liu, Yixin Wan, Wenxuan Wang, Kai-Wei Chang, and Michael R Lyu. 2025b · 2025
Closest in time.
Weidi Luo, Qiming Zhang, Tianyu Lu, Xiaogeng Liu, Yue Zhao, Zhen Xiang, and Chaowei Xiao. 2025 · 2025
Closest in time.
Gauging, enriching and applying geography knowledge in pre-trained language models
Nitin Ramrakhiyani, Vasudeva Varma, Girish Keshav Palshikar, and Sachin Pawar. 2025 · 2025
Closest in time.
Fairgamer: Evaluating biases in the application of large language models to video games
Bingkang Shi, Jen-tse Huang, Guoyi Li, Xiaodan Zhang, and Zhongjiang Yao. 2025 · 2025
Closest in time.
The male ceo and the female assistant: Evaluation and mitigation of gender biases in text-to-image generation of dual subjects
Yixin Wan and Kai-Wei Chang. 2025 · 2025
Closest in time.