Fetching the paper…
Reading the bibliography…
The generative AI revolution in recent years has been spurred by an expansion in compute power and data quantity, which together enable extensive pre-training of powerful text-to-image (T2I) models.
Crowdsourcing with fairness, diversity and budget constraints
Naman Goel and Boi Faltings · 2019
Earlier work this paper cites.
Improving reproducibility of crowdsourcing experiments
Kohta Katsuno, Masaki Matsubara, Chiemi Watanabe, and Atsuyuki Morishima · 2019
Earlier work this paper cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky · 2019
Earlier work this paper cites.
Metrology for AI: From benchmarks to instruments
Chris Welty, Praveen Paritosh, and Lora Aroyo · 2019
Earlier work this paper cites.
Beat the AI: Investigating adversarial human annotation for reading comprehension
Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp · 2020
Earlier work this paper cites.
Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing
Inioluwa Deborah Raji, Andrew Smart, Rebecca N White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes · 2020
Earlier work this paper cites.
Uncovering unknown unknowns in machine learning
Lora Aroyo and Praveen Paritosh · 2021
Earlier work this paper cites.
Multimodal datasets: Misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe · 2021
Earlier work this paper cites.
What will it take to fix benchmarking in natural language understanding?
Samuel Bowman and George Dahl · 2021
Earlier work this paper cites.
Excavating AI: The politics of images in machine learning training sets
Kate Crawford and Trevor Paglen · 2021
Earlier work this paper cites.
The disagreement deconvolution: Bringing machine learning performance metrics in line with reality
Mitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S. Bernstein · 2021
Earlier work this paper cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, et al · 2021
Earlier work this paper cites.
What’s in the box? a preliminary analysis of undesirable content in the common crawl corpus
Alexandra Sasha Luccioni and Joseph D Viviano · 2021
Cited alongside, same era.
DynaSent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela · 2021
Cited alongside, same era.
Zero-shot text-to-image generation, 2021
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
Cited alongside, same era.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela · 2021
Cited alongside, same era.
Dataperf: Benchmarks for data-centric AI development
Mark Mazumder, Colby Banbury, Xiaozhe Yao, Bojan Karlaš, William Gaviria Rojas, Sudnya Diamos, Greg Diamos, Lynn He, Douwe Kiela, David Jurado, et al · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with CLIP latents, 2022
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
Red-teaming the stable diffusion safety filter, 2022
Javier Rando, Daniel Paleka, David Lindner, Lennart Heim, and Florian Tramèr · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding, 2022
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi · 2022
Later among the works it cites.
Dynatask: A framework for creating dynamic AI benchmark tasks
Tristan Thrush, Kushal Tirumala, Anmol Gupta, Max Bartolo, Pedro Rodriguez, Tariq Kane, William Gaviria Rojas, Peter Mattson, Adina Williams, and Douwe Kiela · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guillaume Wenzek, Vishrav Chaudhary, Angela Fan, Sahir Gomez, Naman Goyal, Somya Jain, Douwe Kiela, Tristan Thrush, and Francisco Guzmán · 2021
Cited alongside, same era.
Models in the loop: Aiding crowdworkers with generative annotation assistants
Max Bartolo, Tristan Thrush, Sebastian Riedel, Pontus Stenetorp, Robin Jia, and Douwe Kiela · 2022
Cited alongside, same era.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generative transformers
Jaemin Cho, Abhay Zala, and Mohit Bansal · 2022
Cited alongside, same era.
How Microsoft and Google use AI red teams to “stress test” their systems, 2022
Hayden Field · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Cited alongside, same era.
Handling and presenting harmful text in NLP research
Hannah Kirk, Abeba Birhane, Bertie Vidgen, and Leon Derczynski · 2022
Cited alongside, same era.
Hatemoji: A test suite and adversarially-generated dataset for benchmarking and detecting emoji-based hate
Hannah Kirk, Bertie Vidgen, Paul Röttger, Tristan Thrush, and Scott Hale · 2022
Cited alongside, same era.
Later among the works it cites.
Assessing language model deployment with risk cards
Leon Derczynski, Hannah Rose Kirk, Vidhisha Balachandran, Sachin Kumar, Yulia Tsvetkov, MR Leiser, and Saif Mohammad · 2023
Closest in time.
Midjourney documentation and user guide
Midjourney · 2023
Closest in time.
Auditing large language models: A three-layered approach
Jakob Mökander, Jonas Schuett, Hannah Rose Kirk, and Luciano Floridi · 2023
Closest in time.
DALL·E 2 pre-training mitigations
OpenAI · 2023
Closest in time.
Supporting Human-AI collaboration in auditing LLMs with LLMs
Charvi Rastogi, Marco Tulio Ribeiro, Nicholas King, and Saleema Amershi · 2023
Closest in time.
Announcing the NeurIPS 2021 Datasets and Benchmarks Track | by Neural Information Processing Systems conference | medium
Joaquin Vanschoren and Serena Yeung · 2023
Closest in time.