Fetching the paper…
Reading the bibliography…
Recent progress in generative models has resulted in models that produce both realistic as well as relevant images for most textual inputs.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Generating images from captions with attention
Elman Mansimov, Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, et al · 2017
Earlier work this paper cites.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas · 2017
Earlier work this paper cites.
Indiviuals using the internet
The World Bank · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Earlier work this paper cites.
Women also snowboard: Overcoming bias in captioning models
Lisa Anne Hendricks, Kaylee Burns, Kate Saenko, Trevor Darrell, and Anna Rohrbach · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng · 2019
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, et al · 2019
Earlier work this paper cites.
Dm-gan: Dynamic memory generative adversarial networks for text-to-image synthesis
Minfeng Zhu, Pingbo Pan, Wei Chen, and Yi Yang · 2019
Earlier work this paper cites.
X-lxmert: Paint, caption and answer questions with multi-modal transformers
Jaemin Cho, Jiasen Lu, Dustin Schwenk, Hannaneh Hajishirzi, and Aniruddha Kembhavi · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
Cited alongside, same era.
Measuring social biases in grounded vision and language embeddings
Candace Ross, Boris Katz, and Andrei Barbu · 2020
Cited alongside, same era.
Df-gan: Deep fusion generative adversarial networks for text-to-image synthesis
Ming Tao, Hao Tang, et al · 2020
Cited alongside, same era.
Investigating bias in image classification using model explanations
Schrasing Tong and Lalana Kagal · 2020
Cited alongside, same era.
Understanding and evaluating racial biases in image captioning
Dora Zhao, Angelina Wang, and Olga Russakovsky · 2021
Later among the works it cites.
Dall·e 2: Extending creativity, Jul 2022
2022
Later among the works it cites.
How well can text-to-image generative models understand ethical natural language interventions?
Hritik Bansal, Da Yin, Masoud Monajatipoor, and Kai-Wei Chang · 2022
Later among the works it cites.
Easily accessible text-to-image generation amplifies demographic stereotypes at large scale
Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan · 2022
Later among the works it cites.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generative transformers
Jaemin Cho, Abhay Zala, and Mohit Bansal · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards accuracy-fairness paradox: Adversarial example-based data augmentation for visual debiasing
Yi Zhang and Jitao Sang · 2020
Cited alongside, same era.
Internet/broadband fact sheet
Pew Research Center · 2021
Cited alongside, same era.
An image of society: Gender and racial representation and impact in image search results for occupations
Danaë Metaxa, Michelle A Gan, Su Goh, Jeff Hancock, and James A Landay · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, et al · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, et al · 2021
Cited alongside, same era.
Fair attribute classification through latent space de-biasing
Vikram V Ramaswamy, Sunnie SY Kim, and Olga Russakovsky · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, et al · 2021
Cited alongside, same era.
Later among the works it cites.
Mitigating gender bias in face recognition using the von mises-fisher mixture model
Jean-Rémy Conti, Nathan Noiry, Stephan Clemencon, Vincent Despiegel, and Stéphane Gentric · 2022
Later among the works it cites.
Vector quantized diffusion model for text-to-image synthesis
Shuyang Gu, Dong Chen, et al · 2022
Later among the works it cites.
Underspecification in scene description-to-depiction tasks
Ben Hutchinson, Jason Baldridge, and Vinodkumar Prabhakaran · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, et al · 2022
Later among the works it cites.
The biased artist: Exploiting cultural biases via homoglyphs in text-guided image generation models
Lukas Struppek, Dominik Hintersdorf, and Kristian Kersting · 2022
Later among the works it cites.
Markedness in visual semantic ai
Robert Wolfe and Aylin Caliskan · 2022
Later among the works it cites.
Scaling autoregressive models for content-rich text-to-image generation
Jiahui Yu, Yuanzhong Xu, et al · 2022
Later among the works it cites.
Beyond web-scraping: Crowd-sourcing a geographically diverse image dataset
Vikram V Ramaswamy, Sing Yu Lin, et al · 2023
Closest in time.