I. P. Fellegi and A. B. Sunter, “A theory for record linkage,” Journal of the American Statistical Association , vol. 64, no. 328, pp. 1183–1210, 1969
1969
Earlier work this paper cites.
R. J. A. Little and D. B. Rubin, Statistical Analysis with Missing Data . John Wiley & Sons, 1987
1987
Earlier work this paper cites.
M. Basseville, “Detecting changes in signals and systems—a survey,” Automatica , vol. 24, no. 3, pp. 309–326, 1988
1988
Earlier work this paper cites.
E. Rahm and H. H. Do, “Data cleaning: Problems and current approaches,” IEEE Data Eng. Bull. , vol. 23, no. 4, pp. 3–13, 2000
2000
Earlier work this paper cites.
P. Chapman, J. Clinton, R. Kerber, T. Khabaza, T. Reinartz, C. Shearer, and R. Wirth, “CRISP-DM 1.0 Step-by-step data mining guide,” 2000
2000
Earlier work this paper cites.
Y. Chen, X. S. Zhou, and T. S. Huang, “One-class SVM for learning in image retrieval,” in Proc. International Conf. on Image Processing (ICIP) , 2001, pp. 34–37
2001
Earlier work this paper cites.
L. Breiman, “Random forests,” Machine Learning , vol. 45, no. 1, pp. 5–32, 2001
2001
Earlier work this paper cites.
J. L. Schafer and J. W. Graham, “Missing data: our view of the state of the art.” Psychological Methods , vol. 7, no. 2, p. 147, 2002
2002
Earlier work this paper cites.
E. Eskin, A. Arnold, M. Prerau, L. Portnoy, and S. Stolfo, “A geometric framework for unsupervised anomaly detection,” in Applications of Data Mining in Computer Security . Springer, 2002, pp. 77–101
2002
Earlier work this paper cites.
T. Dasu and T. Johnson, Exploratory Data Mining and Data Cleaning . John Wiley & Sons, 2003, vol. 479
2003
Earlier work this paper cites.
W. Kim, B.-J. Choi, E.-K. Hong, S.-K. Kim, and D. Lee, “A taxonomy of dirty data,” Data Mining and Knowledge Discovery , vol. 7, no. 1, pp. 81–99, 2003
2003
Earlier work this paper cites.
H. Zou and T. Hastie, “Regularization and variable selection via the elastic net,” Journal of the Royal Statistical Society: Series B (Statistical Methodology) , vol. 67, no. 2, pp. 301–320, 2005
2005
Earlier work this paper cites.
A. K. Elmagarmid, P. G. Ipeirotis, and V. S. Verykios, “Duplicate record detection: A survey,” IEEE Transactions on Knowledge and Data Engineering , vol. 19, no. 1, pp. 1–16, 2006
2006
Earlier work this paper cites.
R. K. Pearson, “The problem of disguised missing data,” ACM SIGKDD Explorations Newsletter , vol. 8, no. 1, pp. 83–92, 2006
2006
Earlier work this paper cites.
D. Nadeau and S. Sekine, “A survey of named entity recognition and classification,” Lingvisticae Investigationes , vol. 30, no. 1, pp. 3–26, 2007
2007
Earlier work this paper cites.
A. Poggi, D. Lembo, D. Calvanese, G. De Giacomo, M. Lenzerini, and R. Rosati, “Linking data to ontologies,” J. Data Semantics , vol. 10, pp. 133–173, 2008
2008
Earlier work this paper cites.
F. T. Liu, K. M. Ting, and Z. H. Zhou, “Isolation forest,” in 2008 Eighth IEEE International Conference on Data Mining . IEEE, 2008, pp. 413–422
2008
Earlier work this paper cites.
P. Vassiliadis, “A survey of extract–transform–load technology,” International Journal of Data Warehousing and Mining (IJDWM) , vol. 5, no. 3, pp. 1–27, 2009
2009
Earlier work this paper cites.
V. Chandola, A. Banerjee, and V. Kumar, “Anomaly detection: A survey,” ACM Computing Surveys (CSUR) , vol. 41, no. 3, p. 15, 2009
2009
Earlier work this paper cites.