Fetching the paper…
Reading the bibliography…
Following on recent advances in large language models (LLMs) and subsequent chat models, a new wave of large vision-language models (LVLMs) has emerged.
Racial stereotypes and whites’ political views of blacks in the context of welfare and crime
Mark Peffley, Jon Hurwitz, and Paul M Sniderman. 1997 · 1997
Earlier work this paper cites.
Gender roles and society
Amy M Blackstone. 2003 · 2003
Earlier work this paper cites.
The intersection of gender and race in the labor market
Irene Browne and Joya Misra. 2003 · 2003
Earlier work this paper cites.
Are gender-neutral queries really gender-neutral? mitigating gender bias in image search
Jialu Wang, Yang Liu, and Xin Wang. 2021 · 2008
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara Berg. 2011 · 2011
Earlier work this paper cites.
Social categorization and the perception of social groups
Galen V Bodenhausen, Sonia K Kang, and Destiny Peery. 2012 · 2012
Earlier work this paper cites.
An intersectional analysis of gender and ethnic stereotypes: Testing three hypotheses
Negin Ghavami and Letitia Anne Peplau. 2013 · 2013
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
The myth of racial color blindness: Manifestations, dynamics, and impact
Helen A Neville, Miguel E Gallardo, and Derald Wing Ed Sue. 2016 · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017 · 2017
Earlier work this paper cites.
What racial terms make you cringe?
NYTimes. 2017 · 2017
Earlier work this paper cites.
The measurement of socioeconomic status
J Michael Oakes and Kate E Andrade. 2017 · 2017
Earlier work this paper cites.
Women also snowboard: Overcoming bias in captioning models
Lisa Anne Hendricks, Kaylee Burns, Kate Saenko, Trevor Darrell, and Anna Rohrbach. 2018 · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Earlier work this paper cites.
Exposing and correcting the gender bias in image captioning datasets and models
Shruti Bhargava. 2019 · 2019
Earlier work this paper cites.
Evaluating CLIP: towards characterization of broader capabilities and downstream implications
Sandhini Agarwal, Gretchen Krueger, Jack Clark, Alec Radford, Jong Wook Kim, and Miles Brundage. 2021 · 2021
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. 2021 · 2021
Earlier work this paper cites.
Anti-black racism in academia and what you can do about it
Audrey K Bowden and Cullen R Buie. 2021 · 2021
Earlier work this paper cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. 2021 · 2021
Earlier work this paper cites.
Computer vision and conflicting values: Describing people with automated alt text
Margot Hanley, Solon Barocas, Karen Levy, Shiri Azenkot, and Helen Nissenbaum. 2021 · 2021
Cited alongside, same era.
For black runners, every stride comes with a fear they can’t outrun
Faith Karimi. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Cited alongside, same era.
LAION-400M: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021 · 2021
Cited alongside, same era.
Mitigating gender bias in captioning systems
Ruixiang Tang, Mengnan Du, Yuening Li, Zirui Liu, Na Zou, and Xia Hu. 2021 · 2021
Cited alongside, same era.
InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023 · 2023
Later among the works it cites.
Eva: Exploring the limits of masked visual representation learning at scale
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. 2023 · 2023
Later among the works it cites.
Diversity is not a one-way street: Pilot study on ethical interventions for racial bias in text-to-image systems
Kathleen C. Fraser, Svetlana Kiritchenko, and Isar Nejadgholi. 2023 · 2023
Later among the works it cites.
Vision-language models performing zero-shot tasks exhibit gender-based disparities
Melissa Hall, Laura Gustafson, Aaron Adcock, Ishan Misra, and Candace Ross. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding and evaluating racial biases in image captioning
Dora Zhao, Angelina Wang, and Olga Russakovsky. 2021 · 2021
Cited alongside, same era.
A prompt array keeps the bias away: Debiasing vision-language models with adversarial learning
Hugo Berg, Siobhan Hall, Yash Bhalgat, Hannah Kirk, Aleksandar Shtedritski, and Max Bain. 2022 · 2022
Cited alongside, same era.
Coyo-700m: Image-text pair dataset
Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. 2022 · 2022
Cited alongside, same era.
Gender and racial bias in visual question answering datasets
Yusuke Hirota, Yuta Nakashima, and Noa Garcia. 2022 · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
Underspecification in scene description-to-depiction tasks
Ben Hutchinson, Jason Baldridge, and Vinodkumar Prabhakaran. 2022 · 2022
Cited alongside, same era.
A spontaneous stereotype content model: Taxonomy, properties, and prediction
Gandalf Nicolas, Xuechunzi Bai, and Susan T Fiske. 2022 · 2022
Cited alongside, same era.
Sepehr Janghorbani and Gerard De Melo. 2023 · 2023
Later among the works it cites.
StereoMap: Quantifying the awareness of human-like stereotypes in large language models
Sullam Jeoung, Yubin Ge, and Jana Diesner. 2023 · 2023
Later among the works it cites.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 · 2023
Later among the works it cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 · 2023
Later among the works it cites.
A multi-dimensional study on bias in vision-language models
Gabriele Ruggeri and Debora Nozza. 2023 · 2023
Later among the works it cites.
Dear: Debiasing vision-language models with additive residuals
Ashish Seth, Mayur Hemani, and Chirag Agarwal. 2023 · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Later among the works it cites.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Later among the works it cites.
Contrastive language-vision ai models pretrained on web-scraped multimodal data exhibit sexual objectification bias
Robert Wolfe, Yiwei Yang, Bill Howe, and Aylin Caliskan. 2023 · 2023
Later among the works it cites.
Baize: An open-source chat model with parameter-efficient tuning on self-chat data
Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2023 · 2023
Later among the works it cites.
mPLUG-Owl: Modularization empowers large language models with multimodality
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. 2023 · 2023
Later among the works it cites.
Iti-gen: Inclusive text-to-image generation
Cheng Zhang, Xuanbai Chen, Siqi Chai, Chen Henry Wu, Dmitry Lagun, Thabo Beeler, and Fernando De la Torre. 2023 · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Later among the works it cites.
MiniGPT-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Later among the works it cites.