Fetching the paper…
Reading the bibliography…
Data quality issues have attracted widespread attention due to the negative impacts of dirty data on data mining and machine learning results.
Some methods for classification and analysis of multivariate observations
J. MacQueen et al · 1967
Earlier work this paper cites.
Conditional logit analysis of qualitative choice behavior
D. McFadden et al · 1972
Earlier work this paper cites.
Induction of decision trees
J. R. Quinlan · 1986
Earlier work this paper cites.
Fusion, propagation, and structuring in belief networks
J. Pearl · 1986
Earlier work this paper cites.
Efficient and effective clustering methods for spatial data mining
R. T. Ng and J. Han · 1994
Earlier work this paper cites.
Learning vector quantization
T. Kohonen · 1995
Earlier work this paper cites.
Discriminant adaptive nearest neighbor classification
T. Hastie and R. Tibshirani · 1996
Earlier work this paper cites.
Density-based spatial clustering of applications with noise
M. Ester, H. Kriegel, J. Sander, and X. Xu · 1996
Earlier work this paper cites.
Birch: an efficient data clustering method for very large databases
T. Zhang, R. Ramakrishnan, and M. Livny · 1996
Earlier work this paper cites.
Cure: an efficient clustering algorithm for large databases
S. Guha, R. Rastogi, and K. Shim · 1998
Earlier work this paper cites.
An empirical comparison of supervised learning algorithms
R. Caruana and A. Niculescu-Mizil · 2006
Earlier work this paper cites.
Improving data quality: Consistency and accuracy
G. Cong, W. Fan, F. Geerts, X. Jia, and S. Ma · 2007
Earlier work this paper cites.
An empirical evaluation of supervised learning in high dimensions
R. Caruana, N. Karampatziakis, and A. Yessenalina · 2008
Cited alongside, same era.
Predicting defect-prone software modules using support vector machines
K. O. Elish and M. O. Elish · 2008
Cited alongside, same era.
Capturing missing tuples and missing values
W. Fan and F. Geerts · 2010
Cited alongside, same era.
Crowder: Crowdsourcing entity resolution
J. Wang, T. Kraska, M. J. Franklin, and J. Feng · 2012
Cited alongside, same era.
Foundations of Data Quality Management
W. Fan and F. Geerts · 2012
Cited alongside, same era.
Entity resolution: theory, practice & open challenges
L. Getoor and A. Machanavajjhala · 2012
Cited alongside, same era.
Fastxml: A fast, accurate and stable tree-classifier for extreme multi-label learning
Y. Prabhu and M. Varma · 2014
Later among the works it cites.
Katara: A data cleaning system powered by knowledge bases and crowdsourcing
X. Chu, J. Morcos, I. F. Ilyas, M. Ouzzani, P. Papotti, N. Tang, and Y. Ye · 2015
Later among the works it cites.
Revisiting the impact of classification techniques on the performance of defect prediction models
B. Ghotra, S. McIntosh, and A. E. Hassan · 2015
Later among the works it cites.
Accelerating dynamic time warping clustering with a novel admissible pruning strategy
N. Begum, L. Ulanova, J. Wang, and E. Keogh · 2015
Later among the works it cites.
Heterogeneous network embedding via deep architectures
S. Chang, W. Han, J. Tang, G. Qi, C. C. Aggarwal, and T. S. Huang · 2015
Later among the works it cites.
Optimal action extraction for random forests and boosted trees
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data quality: A survey of data quality dimensions
F. Sidi, P. H. S. Panahy, L. S. Affendey, M. A. Jabar, H. Ibrahim, and A. Mustapha · 2012
Cited alongside, same era.
On the relative trust between inconsistent data and inaccurate constraints
G. Beskales, I. F. Ilyas, L. Golab, and A. Galiullin · 2013
Cited alongside, same era.
Holistic data cleaning: Putting violations into context
X. Chu, I. F. Ilyas, and P. Papotti · 2013
Cited alongside, same era.
Nadeef: a commodity data cleaning system
M. Dallachiesa, A. Ebaid, A. Eldawy, A. Elmagarmid, I. F. Ilyas, M. Ouzzani, and N. Tang · 2013
Cited alongside, same era.
Auto-weka: Combined selection and hyperparameter optimization of classification algorithms
C. Thornton, F. Hutter, H. H. Hoos, and K. Leyton-Brown · 2013
Cited alongside, same era.
Clustering algorithm for incomplete data sets with mixed numeric and categorical attributes
S. Wu, H. Chen, and X. Feng · 2013
Cited alongside, same era.
Z. Cui, W. Chen, Y. He, and Y. Chen · 2015
Later among the works it cites.
Clustering techniques in data mining: A comparison
H. Gulati and P. Singh · 2015
Later among the works it cites.
Extreme multi-label loss functions for recommendation, tagging, ranking & other missing label applications
H. Jain, Y. Prabhu, and M. Varma · 2016
Later among the works it cites.
Facilitating data preprocessing by a generic framework: a proposal for clustering
K. Kirchner, J. Zec, and B. Delibašić · 2016
Later among the works it cites.
Cleaning relations using knowledge bases
S. Hao, N. Tang, G. Li, and J. Li · 2017
Later among the works it cites.
Decomposed normalized maximum likelihood codelength criterion for selecting hierarchical latent variable models
T. Wu, S. Sugawara, and K. Yamanishi · 2017
Later among the works it cites.
An improved overlapping k-means clustering method for medical applications
S. Khanmohammadi, N. Adibeig, and S. Shanehbandy · 2017
Later among the works it cites.