Fetching the paper…
Reading the bibliography…
Datasets have played a foundational role in the advancement of machine learning research.
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni · 1907
Earlier work this paper cites.
“Digital Natives”: How Medical and Indigenous Histories Matter for Big Data
Joanna Radin · 1933
Earlier work this paper cites.
Big Data Is the Answer … But What Is the Question?
Bruno J. Strasser and Paul N. Edwards · 1933
Earlier work this paper cites.
On the issue of roles
Toni Cade Bambara · 1970
Earlier work this paper cites.
Names and faces in the news
T.L. Berg, A.C. Berg, J. Edwards, M. Maire, R. White, Yee-Whye Teh, E. Learned-Miller, and D.A. Forsyth · 2004
Earlier work this paper cites.
On the value of out-of-distribution testing: An example of goodhart’s law
Damien Teney, Kushal Kafle, Robik Shrestha, Ehsan Abbasnejad, Christopher Kanan, and Anton van den Hengel · 2005
Earlier work this paper cites.
80 million tiny images: A large data set for nonparametric object and scene recognition
Antonio Torralba, Rob Fergus, and William T. Freeman · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The unreasonable effectiveness of data
Alon Halevy, Peter Norvig, and Fernando Pereira · 2009
Earlier work this paper cites.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi · 2009
Earlier work this paper cites.
The Pascal Visual Object Classes (VOC) Challenge
Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Critical questions for big data: Provocations for a cultural, technological, and scholarly phenomenon
danah boyd and Kate Crawford · 2012
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2012
Earlier work this paper cites.
A Vast Machine: Computer Models, Climate Data, and the Politics of Global Warming
Paul N. Edwards · 2013
Earlier work this paper cites.
On our best behaviour
Hector J Levesque · 2014
Earlier work this paper cites.
Big data ethics
Neil M Richards and Jonathan H King · 2014
Earlier work this paper cites.
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
The Point of Collection
Mimi Ọnụọha · 2016
Earlier work this paper cites.
Where are human subjects in big data research? the emerging ethics divide
Jacob Metcalf and Kate Crawford · 2016
Earlier work this paper cites.
Seeing through the human reporting bias: Visual classifiers from noisy human-centric labels
Ishan Misra, C. Zitnick, Margaret Mitchell, and Ross Girshick · 2016
Earlier work this paper cites.
Stereotyping and bias in the flickr30k dataset
Emiel van Miltenburg · 2016
Earlier work this paper cites.
Manual Override
Evan Calder Williams · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan · 2017
Earlier work this paper cites.
On the Reuse of Scientific Data
Irene V. Pasquetto, Bernadette M. Randles, and Christine L. Borgman · 2017
Earlier work this paper cites.
Algorithms as culture: Some tactics for the ethnography of algorithmic systems
Nick Seaver · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Earlier work this paper cites.
Men also like shopping: Reducing gender bias amplification using corpus-level constraints
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2017
Earlier work this paper cites.
Do algorithms reveal sexual orientation or just expose our stereotypes?
Blaise Agüera y Arcas, Alexander Todorov, and Margaret Mitchell · 2018
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Earlier work this paper cites.
Women also snowboard: Overcoming bias in captioning models
Kaylee Burns, Lisa Anne Hendricks, Trevor Darrell, and Anna Rohrbach · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman · 2018
Earlier work this paper cites.
Word embeddings quantify 100 years of gender and ethnic stereotypes
Nikhil Garg, Londa Schiebinger, Dan Jurafsky, and James Zou · 2018
Earlier work this paper cites.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford · 2018
Cited alongside, same era.
Gaydar and the fallacy of decontextualized measurement
Andrew Gelman, Greggor Mattson, and Daniel Simpson · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith · 2018
Cited alongside, same era.
The dataset nutrition label: A framework to drive higher data quality standards
Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski · 2018
Cited alongside, same era.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
Divyansh Kaushik and Zachary C. Lipton · 2018
Cited alongside, same era.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Emily M. Bender and Alexander Koller · 2020
Closest in time.
Algorithmic colonization of africa
Abeba Birhane · 2020
Closest in time.
Bringing the people back in: Contesting benchmark machine learning datasets
Emily L. Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, Hilary Nicole, and Morgan Klaus Scheuerman · 2020
Closest in time.
Value-laden disciplinary shifts in machine learning
Ravit Dotan and Smitha Milli · 2020
Closest in time.
Utility is in the eye of the user: A critique of nlp leaderboards
Kawin Ethayarajh and Dan Jurafsky · 2020
Closest in time.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How copyright law can fix artificial intelligence’s implicit bias problem
Amanda Levendowski · 2018
Cited alongside, same era.
Text embeddings contain bias. Here’s why that matters
Ben Packer, M. Mitchell, Mario Guajardo-Céspedes, and Yoni Halpern · 2018
Cited alongside, same era.
Winner’s curse? on pace, progress, and empirical rigor
D. Sculley, Jasper Snoek, Alexander B. Wiltschko, and A. Rahimi · 2018
Cited alongside, same era.
Google’s AI Guru Wants Computers to Think More Like Brains , 2018
Tom Simonite · 2018
Cited alongside, same era.
Towards Standardization of Data Licenses: The Montreal Data License
Misha Benjamin, Paul Gagnon, Negar Rostamzadeh, Chris Pal, Yoshua Bengio, and Alex Shee · 2019
Cited alongside, same era.
Excavating AI: The Politics of Images in Machine Learning Training Sets , 2019
Kate Crawford and Trevor Paglen · 2019
Cited alongside, same era.
Does object recognition work for everyone?
Terrance DeVries, Ishan Misra, Changhan Wang, and Laurens van der Maaten · 2019
Cited alongside, same era.
Closest in time.
Garbage in, garbage out? do machine learning application papers in social computing report where human-labeled training data comes from?
R. Stuart Geiger, Kevin Yu, Yanlai Yang, Mindy Dai, Jie Qiu, Rebekah Tang, and Jenny Huang · 2020
Closest in time.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann · 2020
Closest in time.
Measuring social biases of crowd workers using counterfactual queries
Bhavya Ghai, Q Vera Liao, Yunfeng Zhang, and Klaus Mueller · 2020
Closest in time.
Social biases in nlp models as barriers for persons with disabilities
Ben Hutchinson, Vinodkumar Prabhakaran, Emily Denton, Kellie Webster, Yu Zhong, and Stephen Craig Denuyl · 2020
Closest in time.
Shortcuts: Neural networks love to cheat
Jörn-Henrik Jacobsen, Robert Geirhos, and Claudio Michaelis · 2020
Closest in time.
Lessons from archives: Strategies for collecting sociocultural data in machine learning
Eun Seo Jo and Timnit Gebru · 2020
Closest in time.
Germeval 2020 task 1 on the classification and regression of cognitive and emotional style from text: Companion paper
Dirk Johannßen, Chris Biemann, Steffen Remus, Timo Baumann, and David Sheffer · 2020
Closest in time.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary C Lipton · 2020
Closest in time.
The legality of computer vision datasets
Mehtab Khan and Alex Hanna · 2020
Closest in time.
Adversarial filters of dataset biases
Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew E Peters, Ashish Sabharwal, and Yejin Choi · 2020
Closest in time.
Between subjectivity and imposition: Power dynamics in data annotation for computer vision
Milagros Miceli, Martin Schuessler, and Tianling Yang · 2020
Closest in time.
Decolonial ai: Decolonial theory as sociotechnical foresight in artificial intelligence
Shakir Mohamed, Marie-Therese Png, and William Isaac · 2020
Closest in time.
On lacework: watching an entire machine-learning dataset
Everest Pipkin · 2020
Closest in time.
Large image datasets: A pyrrhic win for computer vision?
Vinay Uday Prabhu and Abeba Birhane · 2020
Closest in time.
The discomfort of death counts: Mourning through the distorted lens of reported covid-19 death data
Inioluwa Deborah Raji · 2020
Closest in time.
Learning machine learning with personal data helps stakeholders ground advocacy arguments in model mechanics
Yim Register and Amy J Ko · 2020
Closest in time.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2020
Closest in time.
How we’ve taught algorithms to see identity: Constructing race and gender in image databases for facial analysis
Morgan Klaus Scheuerman, Kandrea Wade, Caitlin Lustig, and Jed R. Brubaker · 2020
Closest in time.
Targeting the benchmark: On methodology in current natural language processing research
David Schlangen · 2020
Closest in time.
Viktor Schlegel, Goran Nenadic, and Riza Batista-Navarro · 2020
Closest in time.
Robustness to spurious correlations via human annotations
Megha Srivastava, Tatsunori Hashimoto, and Percy Liang · 2020
Closest in time.
The data science life cycle: A disciplined approach to advancing data science as a science
Victoria Stodden · 2020
Closest in time.
Assessing the benchmarking capacity of machine reading comprehension datasets
Saku Sugawara, Pontus Stenetorp, Kentaro Inui, and Akiko Aizawa · 2020
Closest in time.
From imagenet to image classification: Contextualizing progress on benchmarks
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas, and Aleksander Madry · 2020
Closest in time.
Towards fairer datasets: Filtering and balancing the distribution of the people subtree in the imagenet hierarchy
Kaiyu Yang, Klint Qinami, Li Fei-Fei, Jia Deng, and Olga Russakovsky · 2020
Closest in time.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme · 2023
Closest in time.
Best Practices for Computational Science: Software Infrastructure and Environments for Reproducible and Extensible Research
Victoria Stodden and Sheila Miguez · 2049
Closest in time.