Fetching the paper…
Reading the bibliography…
As training datasets become increasingly drawn from unstructured, uncontrolled environments such as the web, researchers and industry practitioners have increasingly relied upon data filtering techniques to "filter out the noise" of web-scraped data.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Making technology masculine: men, women and modern machines in America, 1870-1945
Ruth Oldenziel. 1999 · 1945
Earlier work this paper cites.
Section 107. Limitations on exclusive rights: Fair use
Copyright Law of the United States. 1976 · 1976
Earlier work this paper cites.
When reality monitoring fails: The role of imagination in stereotype maintenance
Morgan P Slusher and Craig A Anderson. 1987 · 1987
Earlier work this paper cites.
The social construction of race
Ian F Haney Lopez. 1995 · 1995
Earlier work this paper cites.
Unpacking hetero-patriarchy: tracing the conflation of sex, gender & (and) sexual orientation to its origins
Francisco Valdes. 1996 · 1996
Earlier work this paper cites.
Breadwinner status and gender ideologies of men and women regarding family roles
Jiping Zuo and Shengming Tang. 2000 · 2000
Earlier work this paper cites.
A survey of user-centered design practice. In Proceedings of the SIGCHI conference on Human factors in computing systems . ACM, Minneapolis, MN, USA, 471–478
Karel Vredenburg, Ji-Ye Mao, Paul W Smith, and Tom Carey. 2002 · 2002
Earlier work this paper cites.
WebInSight: making web images accessible. In Proceedings of the 8th International ACM SIGACCESS Conference on Computers and Accessibility . ACM, Portland, OR, USA, 181–188
Jeffrey P Bigham, Ryan S Kaminsky, Richard E Ladner, Oscar M Danielsson, and Gordon L Hempton. 2006 · 2006
Earlier work this paper cites.
Web content accessibility guidelines (WCAG) 2.0
Ben Caldwell, Michael Cooper, Loretta Guarino Reid, Gregg Vanderheiden, Wendy Chisholm, John Slatin, and Jason White. 2008 · 2008
Earlier work this paper cites.
Using the Census Bureau’s surname list to improve estimates of race/ethnicity and associated disparities
Marc N Elliott, Peter A Morrison, Allen Fremont, Daniel F McCaffrey, Philip Pantoja, and Nicole Lurie. 2009 · 2009
Earlier work this paper cites.
IP geolocation databases: Unreliable?
Ingmar Poese, Steve Uhlig, Mohamed Ali Kaafar, Benoit Donnet, and Bamba Gueye. 2011 · 2011
Earlier work this paper cites.
Fairness through awareness. In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference . ACM, Cambridge, MA, USA, 214–226
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. 2012 · 2012
Earlier work this paper cites.
Gender as performance
Judith Butler. 2013 · 2013
Earlier work this paper cites.
O* NET: The occupational information network
Jonathan D Levine and Frederick L Oswald. 2013 · 2013
Earlier work this paper cites.
langdetect
Nakatani Shuyo. 2014 · 2014
Earlier work this paper cites.
On technological determinism: A typology, scope conditions, and a mechanism
Allan Dafoe. 2015 · 2015
Earlier work this paper cites.
The Chicago face database: A free stimulus set of faces and norming data
Debbie S Ma, Joshua Correll, and Bernd Wittenbrink. 2015 · 2015
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016 · 2016
Earlier work this paper cites.
Is exposure to online content depicting risky behavior related to viewers’ own risky behavior offline?
Dawn Beverley Branley and Judith Covey. 2017 · 2017
Earlier work this paper cites.
Location accuracy of commercial IP address geolocation databases
Dan Komosny, Miroslav Voznak, and Saeed Ur Rehman. 2017 · 2017
Earlier work this paper cites.
Counterfactual fairness
Matt J Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva. 2017 · 2017
Earlier work this paper cites.
Economies of reputation: The case of revenge porn
Ganaele Langlois and Andrea Slane. 2017 · 2017
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conference on Fairness, Accountability, and Transparency . PMLR, New York, NY, USA, 77–91
Joy Buolamwini and Timnit Gebru. 2018 · 2018
Earlier work this paper cites.
Why is my classifier discriminatory?
Irene Chen, Fredrik D Johansson, and David Sontag. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society . ACM, Santa Clara, CA, 67–73
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Earlier work this paper cites.
Caption crawler: Enabling reusable alternative text descriptions using reverse image search. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems . ACM, Montreal, QC, Canada, 1–11
Darren Guinness, Edward Cutrell, and Meredith Ringel Morris. 2018 · 2018
Earlier work this paper cites.
The Art of Cybersecurity: Defense in Depth Strategy for Robust Protection
Arif Ali Mughal. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
The risk of racial bias in hate speech detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . ACL, Florence, Italy, 1668–1678
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A Smith. 2019 · 2019
Earlier work this paper cites.
CCNet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2019 · 2019
Earlier work this paper cites.
Predictive inequity in object detection
Benjamin Wilson, Judy Hoffman, and Jamie Morgenstern. 2019 · 2019
Earlier work this paper cites.
Language (technology) is power: A critical survey of" bias" in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Earlier work this paper cites.
Against scale: Provocations and resistances to scale thinking
Alex Hanna and Tina M Park. 2020 · 2020
Earlier work this paper cites.
Lessons from archives: Strategies for collecting sociocultural data in machine learning. In Proceedings of the 2020 conference on fairness, accountability, and transparency . ACM, Barcelona, Spain, 306–316
Eun Seo Jo and Timnit Gebru. 2020 · 2020
Earlier work this paper cites.
MIT takes down 80 Million Tiny Images data set due to racist and offensive content
Khari Johnson. 2020 · 2020
Earlier work this paper cites.
Detecting child sexual abuse material: A comprehensive survey
Hee-Eun Lee, Tatiana Ermakova, Vasilis Ververis, and Benjamin Fabian. 2020 · 2020
Earlier work this paper cites.
On the accuracy of country-level IP geolocation. In Proceedings of the applied networking research workshop . ACM, Online, 67–73
Ioana Livadariu, Thomas Dreibholz, Anas Saeed Al-Selwi, Haakon Bryhni, Olav Lysne, Steinar Bjørnstad, and Ahmed Elmokashfi. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
Diagnosing gender bias in image recognition systems
Carsten Schwemmer, Carly Knight, Emily D Bello-Pardo, Stan Oklobdzija, Martijn Schoonvelde, and Jeffrey W Lockhart. 2020 · 2020
Earlier work this paper cites.
Contrastive multiview coding. In European Conference on Computer Vision . Springer, Online, 776–794
Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020 · 2020
Earlier work this paper cites.
Fairness-Aware Instrumentation of Preprocessing˜ Pipelines for Machine Learning
Ke Yang, Biao Huang, Julia Stoyanovich, and Sebastian Schelter. 2020 · 2020
Earlier work this paper cites.
Evaluating CLIP: Towards Characterization of Broader Capabilities and Downstream Implications
Sandhini Agarwal, Gretchen Krueger, Jack Clark, Alec Radford, Jong Wook Kim, and Miles Brundage. 2021 · 2021
Earlier work this paper cites.
“You’re so exotic looking”: An intersectional analysis of Asian American and Pacific Islander stereotypes
Sameena Azhar, Antonia RG Alvarez, Anne SJ Farina, and Susan Klumpner. 2021 · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency . ACM, Online, 610–623
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Large image datasets: A pyrrhic win for computer vision?. In 2021 IEEE Winter Conference on Applications of Computer Vision (WACV) . IEEE, Online, 1536–1546
Abeba Birhane and Vinay Uday Prabhu. 2021 · 2021
Cited alongside, same era.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. 2021 · 2021
Cited alongside, same era.
Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipeline. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . ACM, Athens, Greece, 981–993
Sumon Biswas and Hridesh Rajan. 2021 · 2021
Cited alongside, same era.
Automated Data Cleaning Can Hurt Fairness in Machine Learning-based Decision Making. In 2023 IEEE 39th International Conference on Data Engineering (ICDE) . IEEE, Anaheim, CA, USA, 3747–3754
Shubha Guha, Falaah Arif Khan, Julia Stoyanovich, and Sebastian Schelter. 2023 · 2023
Later among the works it cites.
Foundation Models and Fair Use
Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A. Lemley, and Percy Liang. 2023 · 2023
Later among the works it cites.
How AI is being abused to create child sexual abuse imagery
IWF. 2023 · 2023
Later among the works it cites.
Talkin’ ‘Bout AI Generation: Copyright and the Generative-AI Supply Chain
Katherine Lee, A. Feder Cooper, and James Grimmelmann. 2023 · 2023
Later among the works it cites.
word_cloud
Andreas Mueller. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bias and fairness in multimodal machine learning: A case study of automated video interviews. In Proceedings of the 2021 International Conference on Multimodal Interaction . ACM, Montreal, QC, Canada, 268–277
Brandon M Booth, Louis Hickman, Shree Krishna Subburaj, Louis Tay, Sang Eun Woo, and Sidney K D’Mello. 2021 · 2021
Cited alongside, same era.
A list of over 5000 US news domains and their social media accounts
B. Clemm von Hohenberg, E. Menchen-Trevino, A. Casas, and M. Wojcieszak. 2021 · 2021
Cited alongside, same era.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021 · 2021
Cited alongside, same era.
An overview of perceptual hashing
Hany Farid. 2021 · 2021
Cited alongside, same era.
Towards accountability for machine learning datasets: Practices from software engineering and infrastructure. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency . ACM, Online, 560–575
Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell. 2021 · 2021
Cited alongside, same era.
What’s in the box? An analysis of undesirable content in the Common Crawl corpus. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers) . ACL, Bangkok, Thailand, 182–189
Alexandra Sasha Luccioni and Joseph Viviano. 2021 · 2021
Cited alongside, same era.
Why Calling Women’Girls’ Is A Bigger Deal Than You May Think
Susan R Madsen. 2021 · 2021
Cited alongside, same era.
Clipcap: Clip prefix for image captioning
Ron Mokady, Amir Hertz, and Amit H Bermano. 2021 · 2021
Cited alongside, same era.
Gender Biases in Automatic Evaluation Metrics for Image Captioning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 8358–8375
Haoyi Qiu, Zi-Yi Dou, Tianlu Wang, Asli Celikyilmaz, and Nanyun Peng. 2023 · 2023
Later among the works it cites.
Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security . ACM, Copenhagen, Denmark, 3403–3417
Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes, Savvas Zannettou, and Yang Zhang. 2023 · 2023
Later among the works it cites.
Clip for all things zero-shot sketch-based image retrieval, fine-grained or not. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . ACM, Vancouver, Canada, 2765–2775
Aneeshan Sain, Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Subhadeep Koley, Tao Xiang, and Yi-Zhe Song. 2023 · 2023
Later among the works it cites.
Sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society . ACM, Montreal, QC, Canada, 723–741
Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio Garcia, et al · 2023
Later among the works it cites.
Identifying and Eliminating CSAM in Generative ML Training Data and Models
David Thiel. 2023 · 2023
Later among the works it cites.
Generative ML and CSAM: Implications and Mitigations
David Thiel, Melissa Stroebel, and Rebecca Portnoff. 2023 · 2023
Later among the works it cites.
Exploitive, illegal photos of children found in the data that trains some AI
Pranshu Verma and Drew Harwell. 2023 · 2023
Later among the works it cites.
Exploring Legal Approaches to Regulating Nonconsensual Deepfake Pornography
Kaylee Williams. 2023 · 2023
Later among the works it cites.
Contrastive language-vision ai models pretrained on web-scraped multimodal data exhibit sexual objectification bias. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency . ACM, Chicago, IL, USA, 1174–1185
Robert Wolfe, Yiwei Yang, Bill Howe, and Aylin Caliskan. 2023 · 2023
Later among the works it cites.
Demystifying clip data
Hu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang, Russell Howes, Vasu Sharma, Shang-Wen Li, Gargi Ghosh, Luke Zettlemoyer, and Christoph Feichtenhofer. 2023 · 2023
Later among the works it cites.
Stability AI. 2024
2024
Closest in time.
Limits of Algorithmic Fair Use
Jacob Alhadeff, Cooper Cuene, and Max Del Real. 2024 · 2024
Closest in time.
Amazon Rekognition
Amazon. 2024 · 2024
Closest in time.
Ethical Considerations for Responsible Data Curation
Jerone Andrews, Dora Zhao, William Thong, Apostolos Modas, Orestis Papakyriakopoulos, and Alice Xiang. 2024 · 2024
Closest in time.
Fake Photos, Real Harm: AOC and the Fight Against AI Porn
Matthieu Bourel. 2024 · 2024
Closest in time.
Canadian Centre for Child Protection
C3P. 2024 · 2024
Closest in time.
Cloudflare API v4 documentation: Get multiple domain details
Cloudflare. 2024 · 2024
Closest in time.
CC BY 4.0
Creative Commons. 2024 · 2024
Closest in time.
Common Crawl. 2024
2024
Closest in time.
DataComp Tracks
DataComp. 2024 · 2024
Closest in time.
DataComp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al · 2024
Closest in time.
Scaling Laws for Data Filtering – Data Curation cannot be Compute Agnostic
Sachin Goyal, Pratyush Maini, Zachary C. Lipton, Aditi Raghunathan, and J. Zico Kolter. 2024 · 2024
Closest in time.
LAION and the Challenges of Preventing AI-Generated CSAM
Ritwik Gupta. 2024 · 2024
Closest in time.
IP2Location Lite IP-Country IPv6 Database
IP2Location. 2024 · 2024
Closest in time.
Stable bias: Evaluating societal representations in diffusion models
Alexandra Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite. 2024 · 2024
Closest in time.
AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters
Li Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren Klein, and Jesse Dodge. 2024 · 2024
Closest in time.
Tech Companies Promise to Try to Do Something About All the AI CSAM They’re Enabling
Emanuel Maiberg. 2024 · 2024
Closest in time.
PhotoDNA
Microsoft. 2024 · 2024
Closest in time.
Midjourney. 2024
2024
Closest in time.
National Center for Missing and Exploited Children
NCMEC. 2024 · 2024
Closest in time.
Compositional abilities emerge multiplicatively: Exploring diffusion models on a synthetic task
Maya Okawa, Ekdeep S Lubana, Robert Dick, and Hidenori Tanaka. 2024 · 2024
Closest in time.
OpenAI and journalism
OpenAI. 2024 · 2024
Closest in time.
Western Countries 2024
World Population Review. 2024 · 2024
Closest in time.
Here’s How Generative AI Depicts Queer People
Reece Rogers. 2024 · 2024
Closest in time.
I’m still trying to generate an AI Asian man and white woman
Mia Sato and Emillia David. 2024 · 2024
Closest in time.
Can I use images on my website?
Shutterstock. 2024 · 2024
Closest in time.
Teen Girls Confront an Epidemic of Deepfake Nudes in Schools
Natasha Singer. 2024 · 2024
Closest in time.
Trolls have flooded X with graphic Taylor Swift AI fakes
Jess Weatherbed. 2024 · 2024
Closest in time.
The WebAIM Million: An annual accessibility analysis of the top 1,000,000 home pages
WebAIM. 2024 · 2024
Closest in time.