Fetching the paper…
Reading the bibliography…
Datasets are central to training machine learning (ML) models.
More work for mother
Ruth Schwartz Cowan et al · 1983
Earlier work this paper cites.
Introduction to WordNet: An on-line lexical database
George A Miller, Richard Beckwith, Christiane Fellbaum, Derek Gross, and Katherine J Miller. 1990 · 1990
Earlier work this paper cites.
Sorting things out: Classification and its consequences
Geoffrey C Bowker and Susan Leigh Star. 2000 · 2000
Earlier work this paper cites.
Authors Guild v. Google, Inc
[n.d.] · 2007
Earlier work this paper cites.
AV Ex Rel. Vanderhye v. iParadigms, LLC
[n.d.]a · 2007
Earlier work this paper cites.
Perfect 10, Inc. v. Amazon. com, Inc
[n.d.]b · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition . Ieee, 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Digital object identifier (DOI®) system
Norman Paskin. 2010 · 2010
Earlier work this paper cites.
Matters of care in technoscience: Assembling neglected things
Maria Puig de la Bellacasa. 2011 · 2011
Earlier work this paper cites.
Towards tracking semantic change by visual analytics. In Association for Computational Linguistics . 305–310
Christian Rohrdantz, Annette Hautli, Thomas Mayer, Miriam Butt, Daniel Keim, and Frans Plank. 2011 · 2011
Earlier work this paper cites.
A comprehensive survey of retracted articles from the scholarly literature
Michael L Grieneisen and Minghua Zhang. 2012 · 2012
Earlier work this paper cites.
All research outputs should be citable
Mark Hahnel. 2012 · 2012
Earlier work this paper cites.
Raw data is an oxymoron
Lisa Gitelman. 2013 · 2013
Earlier work this paper cites.
Auditing algorithms: Research methods for detecting discrimination on internet platforms
Christian Sandvig, Kevin Hamilton, Karrie Karahalios, and Cedric Langbort. 2014 · 2014
Earlier work this paper cites.
Towards making systems forget with machine unlearning. In 2015 IEEE Symposium on Security and Privacy . IEEE, 463–480
Yinzhi Cao and Junfeng Yang. 2015 · 2015
Earlier work this paper cites.
The politics of care in technoscience
Aryn Martin, Natasha Myers, and Ana Viseu. 2015 · 2015
Earlier work this paper cites.
MegaFace and MF2: Million-Scale Face Recognition
GRAIL University of Washington. 2015 · 2015
Earlier work this paper cites.
Diachronic word embeddings reveal statistical laws of semantic change
William L Hamilton, Jure Leskovec, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
MS-Celeb-1M: Challenge of Recognizing One Million Celebrities in the Real World
Microsoft Research. 2016a · 2016
Earlier work this paper cites.
MS-Celeb-1M: Challenge of Recognizing One Million Celebrities in the Real World
Microsoft Research. 2016b · 2016
Earlier work this paper cites.
The problem with bias: Allocative versus representational harms in machine learning. In 9th Annual Conference of the Special Interest Group for Computing, Information and Society
Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017 · 2017
Earlier work this paper cites.
The CERT guide to coordinated vulnerability disclosure
Allen D Householder, Garret Wassermann, Art Manion, and Chris King. 2017 · 2017
Earlier work this paper cites.
Speed, time, infrastructure
Steven J Jackson. 2017 · 2017
Earlier work this paper cites.
Shreya Shankar, Yoni Halpern, Eric Breck, James Atwood, Jimbo Wilson, and D Sculley. 2017 · 2017
Earlier work this paper cites.
Transgender YouTubers had their videos grabbed to train facial recognition software
James Vincent. 2017 · 2017
Earlier work this paper cites.
Temporal characteristics of retracted articles
Judit Bar-Ilan and Gali Halevi. 2018 · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conference on fairness, accountability and transparency . PMLR, 77–91
Joy Buolamwini and Timnit Gebru. 2018 · 2018
Earlier work this paper cites.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2018 · 2018
Earlier work this paper cites.
Diachronic word embeddings and semantic shifts: a survey
Andrey Kutuzov, Lilja Øvrelid, Terrence Szymanski, and Erik Velldal. 2018 · 2018
Earlier work this paper cites.
How copyright law can fix artificial intelligence’s implicit bias problem
Amanda Levendowski. 2018 · 2018
Earlier work this paper cites.
After innovation, turn to maintenance
Andrew L Russell and Lee Vinsel. 2018 · 2018
Earlier work this paper cites.
Problems caused by semantic drift in wordnet synset construction. In 2019 4th International Conference on Computer Science and Engineering (UBMK) . IEEE, 1–5
Özge Bakay, Özlem Ergelen, and Olcay Taner Yıldız. 2019 · 2019
Earlier work this paper cites.
Towards standardization of data licenses: The montreal data license
Misha Benjamin, Paul Gagnon, Negar Rostamzadeh, Chris Pal, Yoshua Bengio, and Alex Shee. 2019 · 2019
Earlier work this paper cites.
Microsoft Pulls Open Facial Recognition Dataset after Financial Times Investigation
Russel Brandom. 2019 · 2019
Earlier work this paper cites.
Algorithmic fair use
Dan L Burk. 2019 · 2019
Cited alongside, same era.
Excavating AI : The politics of training sets for machine learning
Kate Crawford and Trevor Paglen. 2019 · 2019
Cited alongside, same era.
Better, nicer, clearer, fairer: A critical assessment of the movement for ethical artificial intelligence and machine learning. In Proceedings of the 52nd Hawaii international conference on system sciences
Daniel Greene, Anna Lauren Hoffmann, and Luke Stark. 2019 · 2019
Cited alongside, same era.
MegaPixels: Face Recognition Training Datasets
Adam Harvey and Jules LaPlace. 2019 · 2019
Cited alongside, same era.
How photos of your kids are powering surveillance technology
Kashmir Hill and Aaron Krolik. 2019 · 2019
Cited alongside, same era.
’Nerd’, ’Nonsmoker,’ ’Wrongdoer’ : How Might AI Label You
Cade Metz. 2019 · 2019
Radioactive data: tracing through training. In International Conference on Machine Learning . PMLR, 8326–8335
Alexandre Sablayrolles, Matthijs Douze, Cordelia Schmid, and Hervé Jégou. 2020 · 2020
Later among the works it cites.
Towards probabilistic verification of machine unlearning
David Marco Sommer, Liwei Song, Sameer Wagh, and Prateek Mittal. 2020 · 2020
Later among the works it cites.
80 Million Tiny Images (Dataset Removal Notice)
A. Torralba, R. Fergus, and B. Freeman. 2020 · 2020
Later among the works it cites.
The innovation delusion: How our obsession with the new has disrupted the work that matters most
Lee Vinsel and Andrew L Russell. 2020 · 2020
Later among the works it cites.
Towards fairer datasets: Filtering and balancing the distribution of the people subtree in the imagenet hierarchy. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency . 547–558
Kaiyu Yang, Klint Qinami, Li Fei-Fei, Jia Deng, and Olga Russakovsky. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
MS-Celeb-1M: A Dataset and Benchmark for Large-Scale Face Recognition
Microsoft. 2019 · 2019
Cited alongside, same era.
Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency . 220–229
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019 · 2019
Cited alongside, same era.
Microsoft quietly deletes largest public face recognition data set
Madhumita Murgia. 2019 · 2019
Cited alongside, same era.
Who’s using your face? The ugly truth about facial recognition
Madhumita Murgia and Max Harlow. 2019 · 2019
Cited alongside, same era.
Microsoft Deleted a Massive Facial Recognition Database, But It’s Not Dead
J Pearson. 2019 · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Analyzing information leakage of updates to natural language models. In Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security . 363–375
Santiago Zanella-Béguelin, Lukas Wutschitz, Shruti Tople, Victor Rühle, Andrew Paverd, Olga Ohrimenko, Boris Köpf, and Marc Brockschmidt. 2020 · 2020
Later among the works it cites.
Random erasing data augmentation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 13001–13008
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. 2020 · 2020
Later among the works it cites.
Jack Bandy and Nicholas Vincent. 2021 · 2021
Closest in time.
Pathways to Data: From Plans to Datasets. In 2021 ACM/IEEE Joint Conference on Digital Libraries (JCDL) . IEEE, 254–257
Anastasia Bennett, Will Sutherland, Yubing Tian, Megan Finn, and Amelia Acker. 2021 · 2021
Closest in time.
Machine unlearning. In 2021 IEEE Symposium on Security and Privacy (SP) . IEEE, 141–159
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot. 2021 · 2021
Closest in time.
The algorithm audit: Scoring the algorithms that score us
Shea Brown, Jovana Davidovic, and Ali Hasan. 2021 · 2021
Closest in time.
FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology
Federal Trade Commission. 2021 · 2021
Closest in time.
Op-Ed: Yahoo! Answers is shutting down and taking a record of my teenage self with it
Frances Corry. 2021a · 2021
Closest in time.
Why does a platform die? Diagnosing platform death at Friendster’s end
Frances Corry. 2021b · 2021
Closest in time.
Atlas of AI : Power, Politics, and the Planetary Costs of Artificial Intelligence
Kate Crawford. 2021 · 2021
Closest in time.
On the genealogy of machine learning datasets: A critical history of ImageNet
Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, and Hilary Nicole. 2021 · 2021
Closest in time.
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021 · 2021
Closest in time.
Deleting unethical data sets isn’t good enough
Karen Hao. 2021 · 2021
Closest in time.
Exposing AI
Adam Harvey and Jules LaPlace. 2021 · 2021
Closest in time.
A Practical Guide to Data Privacy Laws by Country [2021]
iSight. 2021 · 2021
Closest in time.
Data Lives: How Data Are Made and Shape Our World
Rob Kitchin. 2021 · 2021
Closest in time.
Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research
Bernard Koch, Emily Denton, Alex Hanna, and Jacob G Foster. 2021 · 2021
Closest in time.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2021 · 2021
Closest in time.
Head Detection Based on DR Feature Extraction Network and Mixed Dilated Convolution Module
Junwen Liu, Yongjun Zhang, Jianbin Xie, Yan Wei, Zewei Wang, and Mengjia Niu. 2021 · 2021
Closest in time.
What’s in the Box? An Analysis of Undesirable Content in the Common Crawl Corpus
Alexandra Sasha Luccioni and Joseph D Viviano. 2021 · 2021
Closest in time.
European privacy activists launch international assault on Clearview AI ’s facial recognition service
D Meyer. 2021 · 2021
Closest in time.
Models of diachronic semantic change using word embeddings
Syrielle Montariol. 2021 · 2021
Closest in time.
Pervasive label errors in test sets destabilize machine learning benchmarks
Curtis G. Northcutt, Anish Athalye, and Jonas Mueller. 2021 · 2021
Closest in time.
Mitigating dataset harms requires stewardship: Lessons from 1000 papers
Kenny Peng, Arunesh Mathur, and Arvind Narayanan. 2021 · 2021
Closest in time.
Google, Microsoft, Amazon, FaceFirst Hit with Biometric Privacy Class Actions Centered on IBM’s ‘Diversity in Faces’ Dataset
C. Rizzi. 2021 · 2021
Closest in time.
“Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. In proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–15
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021 · 2021
Closest in time.
Do datasets have politics? Disciplinary values in computer vision dataset development
Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton. 2021 · 2021
Closest in time.
Demystifying the Draft EU Artificial Intelligence Act—Analysing the good, the bad, and the unclear elements of the proposed approach
Michael Veale and Frederik Zuiderveen Borgesius. 2021 · 2021
Closest in time.
A Study of Face Obfuscation in ImageNet
Kaiyu Yang, Jacqueline Yau, Li Fei-Fei, Jia Deng, and Olga Russakovsky. 2021 · 2021
Closest in time.