Fetching the paper…
Reading the bibliography…
Data exploration and quality analysis is an important yet tedious process in the AI pipeline.
Rapid identification of column heterogeneity. In Sixth International Conference on Data Mining (ICDM’06) . IEEE, 159–170
Bing Tian Dai, Nick Koudas, Beng Chin Ooi, Divesh Srivastava, and Suresh Venkatasubramanian. 2006 · 2006
Earlier work this paper cites.
Handling imbalanced datasets: A review
Sotiris Kotsiantis, Dimitris Kanellopoulos, Panayiotis Pintelas, et al · 2006
Earlier work this paper cites.
Automating string processing in spreadsheets using input-output examples
Sumit Gulwani. 2011 · 2011
Earlier work this paper cites.
Efficient and robust automated machine learning. In Advances in neural information processing systems . 2962–2970
Matthias Feurer, Aaron Klein, Katharina Eggensperger, Jost Springenberg, Manuel Blum, and Frank Hutter. 2015 · 2015
Earlier work this paper cites.
UCI Machine Learning Repository
Dheeru Dua and Casey Graff. 2017 · 2017
Earlier work this paper cites.
Neil D Lawrence. 2017 · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman. 2018 · 2018
Cited alongside, same era.
From theory to practice: A data quality framework for classification tasks
David Camilo Corrales, Agapito Ledezma, and Juan Carlos Corrales. 2018 · 2018
Cited alongside, same era.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2018 · 2018
Cited alongside, same era.
The dataset nutrition label: A framework to drive higher data quality standards
Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski. 2018 · 2018
Cited alongside, same era.
FactSheets: Increasing trust in AI services through supplier’s declarations of conformity
Matthew Arnold, Rachel KE Bellamy, Michael Hind, Stephanie Houde, Sameep Mehta, A Mojsilović, Ravi Nair, K Natesan Ramamurthy, Alexandra Olteanu, David Piorkowski, et al · 2019
The ABC of Data: A Classifying Framework for Data Readiness. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 3–16
Laurens A Castelijns, Yuri Maas, and Joaquin Vanschoren. 2019 · 2019
Later among the works it cites.
Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency . 220–229
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019 · 2019
Later among the works it cites.
Towards an End-to-End Human-Centric Data Cleaning Framework. In Proceedings of the Workshop on Human-In-the-Loop Data Analytics . 1–7
El Kindi Rezig, Mourad Ouzzani, Ahmed K Elmagarmid, Walid G Aref, and Michael Stonebraker. 2019 · 2019
Later among the works it cites.
Mithralabel: Flexible dataset nutritional labels for responsible data science. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management . 2893–2896
Chenkai Sun, Abolfazl Asudeh, HV Jagadish, Bill Howe, and Julia Stoyanovich. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
How to automatically document data with the codebook package to facilitate data reuse
Ruben C Arslan. 2019 · 2019
Cited alongside, same era.
dataMaid: your assistant for data cleaning in R
Anne H Petersen and Claus Thorn Ekstrøm. [n.d.]
Cited in the paper.
Overview and Importance of Data Quality for Machine Learning Tasks. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 3561–3562
Abhinav Jain, Hima Patel, Lokesh Nagalapatti, Nitin Gupta, Sameep Mehta, Shanmukha Guttula, Shashank Mujumdar, Shazia Afzal, Ruhi Sharma Mittal, and Vitobha Munigala. 2020 · 2020
Closest in time.