Fetching the paper…
Reading the bibliography…
The Vision-Language Pre-training (VLP) models like CLIP have gained popularity in recent years.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P.; Lai, A.; Hodosh, M.; and Hockenmaier, J. 2014 · 2014
Earlier work this paper cites.
Deep Learning Face Attributes in the Wild
Liu, Z.; Luo, P.; Wang, X.; and Tang, X. 2015 · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T.; Chang, K.-W.; Zou, J. Y.; Saligrama, V.; and Kalai, A. T. 2016 · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Hardt, M.; Price, E.; and Srebro, N. 2016 · 2016
Earlier work this paper cites.
Age Progression/Regression by Conditional Adversarial Autoencoder
Zhang, Z.; Song, Y.; and Qi, H. 2017 · 2017
Earlier work this paper cites.
Fairness-aware ranking in search & recommendation systems with application to linkedin talent search
Geyik, S. C.; Ambler, S.; and Kenthapadi, K. 2019 · 2019
Earlier work this paper cites.
Uniter: Universal image-text representation learning
Chen, Y.-C.; Li, L.; Yu, L.; El Kholy, A.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020 · 2020
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X.; Yin, X.; Li, C.; Zhang, P.; Hu, X.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; et al. 2020 · 2020
Earlier work this paper cites.
Devlbert: Learning deconfounded visio-linguistic representations
Zhang, S.; Jiang, T.; Wang, T.; Kuang, K.; Zhao, Z.; Zhu, J.; Yu, J.; Yang, H.; and Wu, F. 2020 · 2020
Earlier work this paper cites.
Towards accuracy-fairness paradox: Adversarial example-based data augmentation for visual debiasing
Zhang, Y.; and Sang, J. 2020 · 2020
Cited alongside, same era.
Evaluating clip: towards characterization of broader capabilities and downstream implications
Agarwal, S.; Krueger, G.; Clark, J.; Radford, A.; Kim, J. W.; and Brundage, M. 2021 · 2021
Cited alongside, same era.
ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross-and Intra-modal Knowledge Integration
Cui, Y.; Yu, Z.; Wang, C.; Zhao, Z.; Zhang, J.; Wang, M.; and Yu, J. 2021 · 2021
Cited alongside, same era.
Implicit Stereotypes in Pre-Trained Classifiers
Dehouche, N. 2021 · 2021
Cited alongside, same era.
Clip-adapter: Better vision-language models with feature adapters
Gao, P.; Geng, S.; Zhang, R.; Ma, T.; Fang, R.; Zhang, Y.; Li, H.; and Qiao, Y. 2021 · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C.; Vencu, R.; Beaumont, R.; Kaczmarczyk, R.; Mullis, C.; Katta, A.; Coombes, T.; Jitsev, J.; and Komatsuzaki, A. 2021 · 2021
Later among the works it cites.
How Much Can CLIP Benefit Vision-and-Language Tasks?
Shen, S.; Li, L. H.; Tan, H.; Bansal, M.; Rohrbach, A.; Chang, K.-W.; Yao, Z.; and Keutzer, K. 2021 · 2021
Later among the works it cites.
Image representations learned with unsupervised pre-training contain human-like biases
Steed, R.; and Caliskan, A. 2021 · 2021
Later among the works it cites.
A simple baseline for zero-shot semantic segmentation with pre-trained vision-language model
Xu, M.; Zhang, Z.; Wei, F.; Lin, Y.; Cao, Y.; Hu, H.; and Bai, X. 2021 · 2021
Later among the works it cites.
A prompt array keeps the bias away: Debiasing vision-language models with adversarial learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
FairFace: Face Attribute Dataset for Balanced Race, Gender, and Age for Bias Measurement and Mitigation
Karkkainen, K.; and Joo, J. 2021 · 2021
Cited alongside, same era.
Align before fuse: Vision and language representation learning with momentum distillation
Li, J.; Selvaraju, R.; Gotmare, A.; Joty, S.; Xiong, C.; and Hoi, S. C. H. 2021 · 2021
Cited alongside, same era.
An empirical survey of the effectiveness of debiasing techniques for pre-trained language models
Meade, N.; Poole-Dayan, E.; and Reddy, S. 2021 · 2021
Cited alongside, same era.
Clipcap: Clip prefix for image captioning
Mokady, R.; Hertz, A.; and Bermano, A. H. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Are gender-neutral queries really gender-neutral? mitigating gender bias in image search
Wang, J.; Liu, Y.; and Wang, X. E. 2021a
Cited in the paper.
Assessing Multilingual Fairness in Pre-trained Multimodal Representations
Wang, J.; Liu, Y.; and Wang, X. E. 2021b
Cited in the paper.
Berg, H.; Hall, S. M.; Bhalgat, Y.; Yang, W.; Kirk, H. R.; Shtedritski, A.; and Bain, M. 2022 · 2022
Closest in time.
Enabling Multimodal Generation on CLIP via Vision-Language Knowledge Distillation
Dai, W.; Hou, L.; Shang, L.; Jiang, X.; Liu, Q.; and Fung, P. 2022 · 2022
Closest in time.
Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model
Du, Y.; Wei, F.; Zhang, Z.; Shi, M.; Gao, Y.; and Li, G. 2022 · 2022
Closest in time.
CLIP-GEN: Language-Free Training of a Text-to-Image Generator with CLIP
Wang, Z.; Liu, W.; He, Q.; Wu, X.; and Yi, Z. 2022 · 2022
Closest in time.
Evidence for Hypodescent in Visual Semantic AI
Wolfe, R.; Banaji, M. R.; and Caliskan, A. 2022 · 2022
Closest in time.
Conditional prompt learning for vision-language models
Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2022 · 2022
Closest in time.