Fetching the paper…
Reading the bibliography…
Recently pre-trained multimodal models, such as CLIP, have shown exceptional capabilities towards connecting images and natural language.
Unifying vision-and-language tasks via text generation
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. 2021 · 1942
Earlier work this paper cites.
Learning to scale multilingual representations for vision-language tasks
Andrea Burns, Donghyun Kim, D. Wijaya, Kate Saenko, and Bryan A. Plummer. 2020 · 2004
Earlier work this paper cites.
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. 2012 · 2012
Earlier work this paper cites.
Learning fair representations
R. Zemel, Ledell Yu Wu, Kevin Swersky, T. Pitassi, and C. Dwork. 2013 · 2013
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2015 · 2015
Earlier work this paper cites.
Multi30K: Multilingual English-German image descriptions
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. 2016 · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nathan Srebro. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Fair prediction with disparate impact: A study of bias in recidivism prediction instruments
A. Chouldechova. 2017 · 2017
Earlier work this paper cites.
Learning visual n-grams from web data
Ang Li, A. Jabri, Armand Joulin, and Laurens van der Maaten. 2017 · 2017
Earlier work this paper cites.
Fairness constraints: Mechanisms for fair classification
Muhammad Bilal Zafar, I. Valera, M. Gomez-Rodriguez, and K. Gummadi. 2017 · 2017
Earlier work this paper cites.
Men also like shopping: Reducing gender bias amplification using corpus-level constraints
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2017 · 2017
Earlier work this paper cites.
Women also snowboard: Overcoming bias in captioning models
Kaylee Burns, Lisa Anne Hendricks, Trevor Darrell, and Anna Rohrbach. 2018 · 2018
Earlier work this paper cites.
Translation tutorial: 21 fairness definitions and their politics
Arvind Narayanan. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Towards unsupervised image captioning with shared multimodal embeddings
Iro Laina, C. Rupprecht, and N. Navab. 2019 · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Cited alongside, same era.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
Composing text and image for image retrieval - an empirical odyssey
Nam S. Vo, Lu Jiang, C. Sun, K. Murphy, L. Li, Li Fei-Fei, and James Hays. 2019 · 2019
Gender bias in multilingual embeddings and cross-lingual transfer
Jieyu Zhao, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang, and Ahmed Hassan Awadallah. 2020 · 2020
Later among the works it cites.
Unified vision-language pre-training for image captioning and vqa
L. Zhou, H. Palangi, Lei Zhang, H. Hu, Jason J. Corso, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
Pre-trained multilingual-clip encoders
Fredrik Carlsson and Ariel Ekgren. 2021 · 2021
Closest in time.
How linguistically fair are multilingual pre-trained language models?
Monojit Choudhury and Amit Deshpande. 2021 · 2021
Closest in time.
VirTex: Learning Visual Representations from Textual Annotations
Karan Desai and Justin Johnson. 2021 · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the apparent conflict between individual and group fairness
Reuben Binns. 2020 · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
Large-scale adversarial training for vision-and-language representation learning
Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. 2020 · 2020
Cited alongside, same era.
Type B reflexivization as an unambiguous testbed for multilingual multi-task gender bias
Ana Valeria González, Maria Barrett, Rasmus Hvingelby, Kellie Webster, and Anders Søgaard. 2020 · 2020
Cited alongside, same era.
XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Cited alongside, same era.
Closest in time.
Fairface: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation
Kimmo Karkkainen and Jungseock Joo. 2021 · 2021
Closest in time.
Trends in integration of vision and language research: A survey of tasks, datasets, and methods
Aditya Mogadala, Marimuthu Kalimuthu, and Dietrich Klakow. 2021 · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Closest in time.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Closest in time.
Measuring social biases in grounded vision and language embeddings
Candace Ross, B. Katz, and Andrei Barbu. 2021 · 2021
Closest in time.
Worst of both worlds: Biases compound in pre-trained vision-and-language models
Tejas Srinivasan and Yonatan Bisk. 2021 · 2021
Closest in time.
Mitigating gender bias in captioning systems
Ruixiang Tang, Mengnan Du, Yuening Li, Zirui Liu, Na Zou, and Xia Hu. 2021 · 2021
Closest in time.
Are gender-neutral queries really gender-neutral? mitigating gender bias in image search
Jialu Wang, Yang Liu, and Xin Eric Wang. 2021 · 2021
Closest in time.
Ernie-vil: Knowledge enhanced vision-language representations through scene graphs
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2021 · 2021
Closest in time.
The rich get richer: Disparate impact of semi-supervised learning
Zhaowei Zhu, Tianyi Luo, and Yang Liu. 2022 · 2022
Closest in time.