Fetching the paper…
Reading the bibliography…
The quality of training data has a huge impact on the efficiency, accuracy and complexity of machine learning tasks.
Hawkins, D. M.: Identification of Outliers. Chapman and Hall, London – New York 1980, 188 S., £ 14, 50
G. Enderlein. 1987 · 1987
Earlier work this paper cites.
Identifying mislabeled training data
Carla E Brodley and Mark A Friedl. 1999 · 1999
Earlier work this paper cites.
LOF: Identifying Density-Based Local Outliers
Markus Breunig, Hans-Peter Kriegel, Raymond Ng, and Joerg Sander. 2000 · 2000
Earlier work this paper cites.
An empirical study of the effect of outliers on the misclassification error rate
E. Acuña and C. Rodríguez. 2005 · 2005
Earlier work this paper cites.
Shazia Afzal, Rajmohan C, Manish Kesarwani, Sameep Mehta, and Hima Patel. 2020 · 2010
Earlier work this paper cites.
KEEL Data-Mining Software Tool: Data Set Repository, Integration of Algorithms and Experimental Analysis Framework
Jesus Alcala-Fdez, Alberto Fernández, Julián Luengo, J. Derrac, S Garc’ia, Luciano Sanchez, and Francisco Herrera. 2010 · 2010
Cited alongside, same era.
Comparison of Values of Pearson’s and Spearman’s Correlation Coefficients on the Same Sets of Data
Jan Hauke and Tomasz Kossowski. 2011 · 2011
Cited alongside, same era.
The Precise Effect of Multicollinearity on Classification Prediction
Mary Lieberman and John Morris. 2014 · 2014
Cited alongside, same era.
Rachel KE Bellamy, Kuntal Dey, Michael Hind, Samuel C Hoffman, Stephanie Houde, Kalapriya Kannan, Pranay Lohia, Jacquelyn Martino, Sameep Mehta, Aleksandra Mojsilovic, et al · 2018
Cited alongside, same era.
Confident learning: Estimating uncertainty in dataset labels
Curtis G Northcutt, Lu Jiang, and Isaac L Chuang. 2019 · 2019
Addressing the Overlapping Data Problem in Classification Using the One-vs-One Decomposition Strategy
J. A. Sáez, M. Galar, and B. Krawczyk. 2019 · 2019
Later among the works it cites.
Overview and Importance of Data Quality for Machine Learning Tasks. In KDD
Abhinav Jain, Hima Patel, Lokesh Nagalapatti, Nitin Gupta, Sameep Mehta, Shanmukha Guttula, Shashank Mujumdar, Shazia Afzal, Ruhi Sharma Mittal, and Vitobha Munigala. 2020 · 2020
Later among the works it cites.
An ADMM based framework for automl pipeline configuration. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 4892–4899
Sijia Liu, Parikshit Ram, Deepak Vijaykeerthy, Djallel Bouneffouf, Gregory Bramble, Horst Samulowitz, Dakuo Wang, Andrew Conn, and Alexander Gray. 2020 · 2020
Later among the works it cites.
Comparison of Outlier Detection Techniques for Structured Data
Amulya Agarwal and Nitin Gupta. 2021 · 2021
Closest in time.
1st International Workshop on Data Assessment and Readiness for AI.. In PAKDD (Workshops)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
IBM API Hub
[n.d.]a
Cited in the paper.
IBM Data Quality for AI APIs
[n.d.]
Cited in the paper.
IBM Learning Path
[n.d.]b
Cited in the paper.
Kaggle Repository
[n.d.]
Cited in the paper.
UCI Machine Learning Repository
Dheeru Dua and Casey Graff. 2017a
Cited in the paper.
UCI Machine Learning Repository
Dheeru Dua and Casey Graff. 2017b
Cited in the paper.
Data Quality for Machine Learning Tasks. In KDD
Nitin Gupta, Shashank Mujumdar, Hima Patel, Satoshi Masuda, Naveen Panwar, Sambaran Bandyopadhyay, Sameep Mehta, Shanmukha Guttula, Shazia Afzal, Ruhi Sharma Mittal, and Vitobha Munigala. 2021a
Cited in the paper.
Bortik Bandyopadhyay, Sambaran Bandyopadhyay, Srikanta Bedathur, Nitin Gupta, Sameep Mehta, Shashank Mujumdar, Srinivasan Parthasarathy, and Hima Patel. 2021 · 2021
Closest in time.