Fetching the paper…
Reading the bibliography…
Data quality affects machine learning (ML) model performances, and data scientists spend considerable amount of time on data cleaning before model training.
Teoria statistica delle classi e calcolo delle probabilità
C. E. Bonferroni · 1936
Earlier work this paper cites.
Simplifying decision trees
J. R. Quinlan · 1987
Earlier work this paper cites.
Controlling the false discovery rate: a practical and powerful approach to multiple testing
Y. Benjamini and Y. Hochberg · 1995
Earlier work this paper cites.
Consistent query answers in inconsistent databases
M. Arenas, L. Bertossi, and J. Chomicki · 1999
Earlier work this paper cites.
Data cleaning: Problems and current approaches
E. Rahm and H. H. Do · 2000
Earlier work this paper cites.
Evaluating noise correction
C. M. Teng · 2000
Earlier work this paper cites.
The elements of statistical learning
J. Friedman, T. Hastie, and R. Tibshirani · 2001
Earlier work this paper cites.
Building classification trees using the total uncertainty criterion
J. Abellán and S. Moral · 2003
Earlier work this paper cites.
Conditional functional dependencies for data cleaning
P. Bohannon, W. Fan, F. Geerts, X. Jia, and A. Kementsietsidis · 2007
Earlier work this paper cites.
Duplicate record detection: A survey
A. K. Elmagarmid, P. G. Ipeirotis, and V. S. Verykios · 2007
Earlier work this paper cites.
Quantitative data cleaning for large databases
J. M. Hellerstein · 2008
Earlier work this paper cites.
Handbook of biological statistics
J. H. McDonald · 2009
Earlier work this paper cites.
Comparing boosting and bagging techniques with noisy and imbalanced data
T. M. Khoshgoftaar, J. Van Hulse, and A. Napolitano · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Detection of outliers and influential observations in binary logistic regression: An empirical study
S. K. Sarkar, H. Midi, and S. Rana · 2011
Earlier work this paper cites.
Foundations of data quality management
W. Fan and F. Geerts · 2012
Earlier work this paper cites.
Simultaneous statistical inference
G. Rupert Jr et al · 2012
Earlier work this paper cites.
Outlier Analysis
C. C. Aggarwal · 2013
Cited alongside, same era.
Discovering denial constraints
X. Chu, I. F. Ilyas, and P. Papotti · 2013
Cited alongside, same era.
Holistic data cleaning: Putting violations into context
X. Chu, I. F. Ilyas, and P. Papotti · 2013
Cited alongside, same era.
Using OpenRefine
R. Verborgh and M. De Wilde · 2013
Cited alongside, same era.
Classification in the presence of label noise: a survey
B. Frénay and M. Verleysen · 2014
Cited alongside, same era.
A sample-and-clean framework for fast and accurate query processing on dirty data
J. Wang, S. Krishnan, M. J. Franklin, K. Goldberg, T. Kraska, and T. Milo · 2014
Cited alongside, same era.
Messing up with bart: error generation for evaluating data-cleaning algorithms
Boostclean: Automated error detection and repair for machine learning
S. Krishnan, M. J. Franklin, K. Goldberg, and E. Wu · 2017
Later among the works it cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu · 2017
Later among the works it cites.
Learning with confident examples: Rank pruning for robust classification with noisy labels
C. G. Northcutt, T. Wu, and I. L. Chuang · 2017
Later among the works it cites.
Data management challenges in production machine learning
N. Polyzotis, S. Roy, S. E. Whang, and M. Zinkevich · 2017
Later among the works it cites.
Holoclean: Holistic data repairs with probabilistic inference
T. Rekatsinas, X. Chu, I. F. Ilyas, and C. Ré · 2017
Later among the works it cites.
Deep learning is robust to massive label noise
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. C. Arocena, B. Glavic, G. Mecca, R. J. Miller, P. Papotti, and D. Santoro · 2015
Cited alongside, same era.
Katara: A data cleaning system powered by knowledge bases and crowdsourcing
X. Chu, J. Morcos, I. F. Ilyas, M. Ouzzani, P. Papotti, N. Tang, and Y. Ye · 2015
Cited alongside, same era.
Data preprocessing in data mining
S. García, J. Luengo, and F. Herrera · 2015
Cited alongside, same era.
Trends in cleaning relational data: Consistency and deduplication
I. F. Ilyas, X. Chu, et al · 2015
Cited alongside, same era.
Staleness-aware async-sgd for distributed deep learning
W. Zhang, S. Gupta, X. Lian, and J. Liu · 2015
Cited alongside, same era.
https://www.forbes.com/sites/gilpress/2016/03/23/data-preparation-most-time-consuming-least-enjoyable-data-science-task-survey-says/
Cleaning big data: Most time-consuming, least enjoyable data science task · 2016
Cited alongside, same era.
D. Rolnick, A. Veit, S. Belongie, and N. Shavit · 2017
Later among the works it cites.
Zipml: Training linear models with end-to-end low precision, and a little bit of deep learning
H. Zhang, J. Li, K. Kara, D. Alistarh, J. Liu, and C. Zhang · 2017
Later among the works it cites.
Controlling false discoveries during interactive data exploration
Z. Zhao, L. De Stefani, E. Zgraggen, C. Binnig, E. Upfal, and T. Kraska · 2017
Later among the works it cites.
Distributed representations of tuples for entity resolution
M. Ebraheem, S. Thirumuruganathan, S. Joty, M. Ouzzani, and N. Tang · 2018
Later among the works it cites.
Transform-data-by-example (tde): an extensible search engine for data transformations
Y. He, X. Chu, K. Ganjam, Y. Zheng, V. Narasayya, and S. Chaudhuri · 2018
Later among the works it cites.
Benchmarking neural network robustness to common corruptions and surface variations
D. Hendrycks and T. G. Dietterich · 2018
Later among the works it cites.
Deep learning for entity matching: A design space exploration
S. Mudgal, H. Li, T. Rekatsinas, A. Doan, Y. Park, G. Krishnan, R. Deep, E. Arcaute, and V. Raghavendra · 2018
Later among the works it cites.
Holodetect: Few-shot learning for error detection
A. Heidari, J. McGrath, I. F. Ilyas, and T. Rekatsinas · 2019
Closest in time.
What to expect of classifiers? reasoning about logistic regression with missing features
P. Khosravi, Y. Liang, Y. Choi, and G. V. d. Broeck · 2019
Closest in time.
CleanML Technical Report
P. Li, X. Rao, J. Blase, Y. Zhang, X. Chu, and C. Zhang · 2019
Closest in time.
B. Karlaš, P. Li, R. Wu, N. M. Gürel, X. Chu, W. Wu, and C. Zhang · 2020
Closest in time.
Zeroer: Entity resolution using zero labeled examples
R. Wu, S. Chaba, S. Sawlani, X. Chu, and S. Thirumuruganathan · 2020
Closest in time.