Fetching the paper…
Reading the bibliography…
With significant advances in generative AI, new technologies are rapidly being deployed with generative components.
Normalized (pointwise) mutual information in collocation extraction
Gerlof Bouma · 2009
Earlier work this paper cites.
Auto-encoder variational inference
Diederik P. Kingma and Max Welling · 2013
Earlier work this paper cites.
Generative adversarial networks
Ian Goodfellow, Julien Pouget-Abadie, Arthur Courville, Mehdi Mirza, and Thomas Yosinski · 2014
Earlier work this paper cites.
Variational inference for generative models
Danilo J. Rezende, Diederik P. Kingma, and Max Welling · 2014
Earlier work this paper cites.
The virtues of moderation
James Grimmelmann · 2015
Earlier work this paper cites.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, Eric Price, and Nati Srebro · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on twitter
Zeerak Waseem and Dirk Hovy · 2016
Earlier work this paper cites.
Data decisions and theoretical implications when adversarially learning fair representations
Alex Beutel, Jilin Chen, Zhe Zhao, and Ed H Chi · 2017
Earlier work this paper cites.
Like trainer, like bot? inheritance of bias in algorithmic content moderation
Reuben Binns, Michael Veale, Max Van Kleek, and Nigel Shadbolt · 2017
Earlier work this paper cites.
Algorithmic decision making and the cost of fairness
Sam Corbett-Davies, Emma Pierson, Avi Feller, Sharad Goel, and Aziz Huq · 2017
Earlier work this paper cites.
The platform is the message
James Grimmelmann · 2017
Earlier work this paper cites.
Improved variational auto-encoders
Guoyou Huang, Junyi Liu, Adelbert van den Oord, and Diederik P. Kingma · 2017
Earlier work this paper cites.
A survey on hate speech detection using natural language processing
Anna Schmidt and Michael Wiegand · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Sandeep Shyam, Sameer Narang, and et al · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional representations from unlabeled text
Jacob Devlin, Ming-Wei Chang, Kevin Lee, and Alec Radford · 2018
Earlier work this paper cites.
Internet platforms: Observations on speech, danger, and money
Daphne Keller · 2018
Earlier work this paper cites.
Fairness definitions explained
Sahil Verma and Julia Rubin · 2018
Earlier work this paper cites.
Pareto-efficient fairness for skewed subgroup data
Ananth Balashankar, Alyssa Lees, Chris Welty, and Lakshminarayanan Subramanian · 2019
Earlier work this paper cites.
Fairness and Machine Learning
Solon Barocas, Moritz Hardt, and Arvind Narayanan · 2019
Earlier work this paper cites.
Race After Technology: Abolitionist Tools for the New Jim Code
Ruha Benjamin · 2019
Earlier work this paper cites.
Identifying and reducing gender bias in word-level language models
Shikha Bordia and Samuel R. Bowman · 2019
Earlier work this paper cites.
Does object recognition work for everyone?, 2019
Terrance DeVries, Ishan Misra, Changhan Wang, and Laurens van der Maaten · 2019
Earlier work this paper cites.
Counterfactual fairness in text classification through robustness
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H Chi, and Alex Beutel · 2019
Earlier work this paper cites.
Reducing gender bias in word-level language models with a gender-equalizing loss function
Yusu Qian, Urwa Muaz, Ben Zhang, and Jae Won Hyun · 2019
Earlier work this paper cites.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith · 2019
Cited alongside, same era.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith · 2019
Cited alongside, same era.
Challenges and frontiers in abusive content detection
Bertie Vidgen, Alex Harris, Dong Nguyen, Rebekah Tromble, Scott Hale, and Helen Margetts · 2019
Cited alongside, same era.
Content moderation technologies: Applying human rights standards to protect freedom of expression
Thiago Dias Oliva · 2020
Cited alongside, same era.
Content moderation, ai, and the question of scale
Tarleton Gillespie · 2020
Cited alongside, same era.
Algorithmic content moderation: Technical and political challenges in the automation of platform governance
Robert Gorwa, Reuben Binns, and Christian Katzenbach · 2020
Stable diffusion: A scalable and controllable generative model
Xuan Zhang, Xiang Wang, Xin Liu, and et al · 2021
Later among the works it cites.
Anthropic: A framework for ai safety and alignment
Milad Aghajanian, Andrew Berensmeier, Sam Benesty, and et al · 2022
Later among the works it cites.
Easily accessible text-to-image generation amplifies demographic stereotypes at large scale
Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan · 2022
Later among the works it cites.
The expertise problem: Learning from specialized feedback, 2022
Oliver Daniels-Koch and Rachel Freedman · 2022
Later among the works it cites.
Do not recommend? reduction as a form of content moderation
Tarleton Gillespie · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ai paradigms and ai safety: mapping artefacts and techniques to safety issues, 2020
Jose Hernández-Orallo, Fernando Martínez-Plumed, Shahar Avin, Jessica Whittlestone, et al · 2020
Cited alongside, same era.
Reducing sentiment bias in language models via counterfactual evaluation
Po-Sen Huang, Huan Zhang, Ray Jiang, Robert Stanforth, Johannes Welbl, Jack Rae, Vishal Maini, Dani Yogatama, and Pushmeet Kohli · 2020
Cited alongside, same era.
A scalable approach to reducing gender bias in google translate
Melvin Johnson · 2020
Cited alongside, same era.
Artificial intelligence in digital media: The era of deepfakes
Stamatis Karnouskos · 2020
Cited alongside, same era.
The hateful memes challenge: Detecting hate speech in multimodal memes
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine · 2020
Cited alongside, same era.
Recipes for safety in open-domain chatbots
Margaret Li, Jason Boureau, Emily Weston, and Da Ju Dinan, Jing Xu · 2020
Cited alongside, same era.
Later among the works it cites.
Something that they never said: Multimodal disinformation and source vividness in understanding the power of ai-enabled deepfake news
Jiyoung Lee and Soo Yun Shin · 2022
Later among the works it cites.
Codegpt: A generative pre-trained transformer for code
Junyi Liu, Adelbert van den Oord, and Ashish Vaswani · 2022
Later among the works it cites.
Chatgpt: A large language model for dialogue generation
Junyi Liu, Xiang Wang, Xin Liu, and et al · 2022
Later among the works it cites.
Cultural incongruencies in artificial intelligence, 2022
Vinodkumar Prabhakaran, Rida Qadri, and Ben Hutchinson · 2022
Later among the works it cites.
Red-teaming the stable diffusion safety filter, 2022
Javier Rando, Daniel Paleka, David Lindner, Lennart Heim, and Florian Tramèr · 2022
Later among the works it cites.
Very neat trick to tease this out
rzhang88@ · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al · 2022
Later among the works it cites.
Identifying sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction, 2022
Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla, Jess Gallegos, Andrew Smart, Emilio Garcia, and Gurleen Virk · 2022
Later among the works it cites.
The recent dall-e 2 update made changes to increase the diversity of generated images
TylerGlaiel@ · 2022
Later among the works it cites.
Measuring representational harms in image captioning
Angelina Wang, Solon Barocas, Kristen Laird, and Hanna Wallach · 2022
Later among the works it cites.
Taxonomy of risks posed by language models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2022
Later among the works it cites.
Dataset publishing language resource, 2023
Google Developers · 2023
Closest in time.
Regulating chatgpt and other large generative ai models, 2023
Philipp Hacker, Andreas Engel, and Marco Mauer · 2023
Closest in time.
Hate on display: Hate symbols database, 2023
Anti Defamation League · 2023
Closest in time.
Towards globally responsible and human-centered text-to-image evaluations, 2023
Rida Qadri, Emily Denton, and Renee Shelby · 2023
Closest in time.
The gradient of generative ai release: Methods and considerations, 2023
Irene Solaiman · 2023
Closest in time.
League of arab states, 2023
European Union · 2023
Closest in time.
Toward general design principles for generative ai applications, 2023
Justin D. Weisz, Michael Muller, Jessica He, and Stephanie Houde · 2023
Closest in time.