Fetching the paper…
Reading the bibliography…
We present Public Domain 12M (PD12M), a dataset of 12.4 million high-quality public domain and CC0-licensed images with synthetic captions, designed for training text-to-image models.
Dataset shift in machine learning
Joaquin Quionero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence · 2009
Earlier work this paper cites.
Unbiased look at dataset bias
Antonio Torralba and Alexei A Efros · 2011
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
The fair guiding principles for scientific data management and stewardship
Mark D Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Appleton, et al · 2016
Earlier work this paper cites.
Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
Detecting and correcting for label shift with black box predictors
Zachary Lipton, Yu-Xiang Wang, and Alex Smola · 2018
Earlier work this paper cites.
Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science
Emily M Bender and Batya Friedman · 2018
Earlier work this paper cites.
The dataset nutrition label: A framework to drive higher data quality standards
Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski · 2018
Earlier work this paper cites.
Artificial intelligence faces reproducibility crisis
Matthew Hutson · 2018
Earlier work this paper cites.
Web scraping with Python: Collecting more data from the modern web
Ryan Mitchell · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric, 2018
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
Failing loudly: An empirical study of methods for detecting dataset shift
Stephan Rabanser, Stephan Günnemann, and Zachary Lipton · 2019
Earlier work this paper cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Earlier work this paper cites.
Directive (EU) 2019/790 on copyright and related rights in the Digital Single Market
European Parliament and Council of the European Union · 2019
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song · 2019
Earlier work this paper cites.
New digital worlds: Postcolonial digital humanities in theory, praxis, and pedagogy
Roopika Risam · 2019
Earlier work this paper cites.
Large image datasets: A pyrrhic win for computer vision?
Vinay Uday Prabhu and Abeba Birhane · 2020
Earlier work this paper cites.
REVISE: A tool for measuring and mitigating bias in visual datasets
Angelina Wang, Arvind Narayanan, and Olga Russakovsky · 2020
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber · 2020
Earlier work this paper cites.
Strategies for collecting sociocultural data in machine learning
Eun Seo Jo and Timnit Gebru · 2020
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Romain Vencu, Ronan Beaumont, Robert Klinker, Mehdi Sajjadi, Aakash Doomasia, and Perick Ogayo · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, et al · 2021
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe · 2021
Earlier work this paper cites.
Data and its (dis)contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna · 2021
Earlier work this paper cites.
Excavating AI: The politics of images in machine learning training sets
Kate Crawford and Trevor Paglen · 2021
Earlier work this paper cites.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford · 2021
Earlier work this paper cites.
LAION-NSFW: A Robust Model for NSFW Detection
Ethan Schonfeld, Tobias Weyand, Kate Saenko, and Rogerio Feris · 2021
Cited alongside, same era.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfa Erlingsson, Alina Oprea, and Colin Raffel · 2021
Cited alongside, same era.
Improving reproducibility in machine learning research
Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha, et al · 2021
Cited alongside, same era.
Systematic review of reproducibility research in natural language processing
Anya Belz, Shubham Agarwal, Anastasia Shimorina, and Ehud Reiter · 2021
Cited alongside, same era.
Accounting for variance in machine learning benchmarks
Xavier Bouthillier, Pierre Delaunay, Mirko Bronzi, et al · 2021
Cited alongside, same era.
Florence-2: Advancing a unified representation for a variety of vision tasks
Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, and Lu Yuan · 2023
Later among the works it cites.
Declaration of David Holz in Support of Midjourney’s Opposition to Plaintiffs’ Motion for Preliminary Injunction, 2023
David Holz · 2023
Later among the works it cites.
An AI Scraping Tool Is Overwhelming Websites With Traffic
Joseph Cox · 2023
Later among the works it cites.
Mapping the Impact of Share Alike/Copyleft Licensing on Machine Learning
Kacper Szkalej and Martin Senftleben · 2024
Closest in time.
Poisoning web-scale training datasets is practical, 2024
Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Vencu, Ronan Beaumont, Robert Klinker, Clayton Lambdalabs, Ksc Hong, Jin Yun Kim, Seo Paquet, Samuel Breitenfeld, Byeongheup Oh, et al · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Public Data Commons
Alek Tarkowski, Paul Keller, Francesco Vogelezang, and Jan Zygmuntowski · 2022
Cited alongside, same era.
Public Data Commons: A Framework for Open Data Infrastructure
Paul Keller and Alek Tarkowski · 2022
Cited alongside, same era.
SSCD: Self-Supervised Copy Detection
Gabriel Pizzi, Mathilde Caron, Thomas Séjourne, Matthijs Douze, Julien Mairal, and Herve Jegou · 2022
Cited alongside, same era.
Jianfeng Gao, Xizhou Chen, Yaorong Li, et al · 2024
Closest in time.
Mitsua Diffusion One: A High-Quality Text-to-Image Generation Model
Mitsua Research · 2024
Closest in time.
Megalith-10m: A Dataset of Public Domain Photographs
DrawThings.ai · 2024
Closest in time.
Digital Public Infrastructure and Public Value: What Is "Public" about DPI?
David Eaves, Marianna Mazzucato, and Beatriz Vasconcellos · 2024
Closest in time.
The Open Source AI Definition
Open Source Initiative · 2024
Closest in time.
Open GLAM Survey: An Ongoing Survey of Open Access Policies, 2024
Douglas McCarthy and Andrea Wallace · 2024
Closest in time.
AI crawlers need to be more respectful
Eric Holscher · 2024
Closest in time.
Wikimedia Commons
Wikimedia Foundation · 2024
Closest in time.
Do Not Train Registry
Spawning AI · 2024
Closest in time.
Conceptual Captions Dataset
Hugging Face · 2024
Closest in time.
CLIP: Contrastive Language-Image Pre-training
OpenAI · 2024
Closest in time.
Sana: Efficient high-resolution image synthesis with linear diffusion transformers, 2024
Enze Xie, Junsong Chen, Junyu Chen, Han Cai, Haotian Tang, Yujun Lin, Zhekai Zhang, Muyang Li, Ligeng Zhu, Yao Lu, and Song Han · 2024
Closest in time.
Pixart- σ \sigma : Weak-to-strong training of diffusion transformer for 4k text-to-image generation, 2024
Junsong Chen, Chongjian Ge, Enze Xie, Yue Wu, Lewei Yao, Xiaozhe Ren, Zhongdao Wang, Ping Luo, Huchuan Lu, and Zhenguo Li · 2024
Closest in time.
List of ethnic slurs
Wikipedia contributors · 2024
Closest in time.
The dumb reason your fancy Computer Vision app isn’t working: Exif Orientation
Adam Geitgey · 2024
Closest in time.
Takedown Policy for Source.Plus
Spawning AI · 2024
Closest in time.
Metadata & Attribution Correction Policy for Source.Plus
Spawning AI · 2024
Closest in time.
Persistent pre-training poisoning of llms, 2024
Yiming Zhang, Javier Rando, Ivan Evtimov, Jianfeng Chi, Eric Michael Smith, Nicholas Carlini, Florian Tramèr, and Daphne Ippolito · 2024
Closest in time.
When online content disappears - 2024
Pew Research Center · 2024
Closest in time.
Flickr API Terms of Use
Flickr · 2024
Closest in time.