Fetching the paper…
Reading the bibliography…
Modern artificial intelligence (AI) applications require large quantities of training and test data.
A General Coefficient of Similarity and Some of its Properties
J. C. Gower. 1971 · 1971
Earlier work this paper cites.
Hedonic Housing Prices and the Demand for Clean Air
D. Harrison and D. L. Rubinfeld. 1978 · 1978
Earlier work this paper cites.
Classification and Regression Trees
Leo Breiman, J. H. Friedman, R. A. Olshen, and C. J. Stone. 1984 · 1984
Earlier work this paper cites.
Efficient Algorithms for Agglomerative Hierarchical Clustering Methods
W. H. E. Day and H. Edelsbrunner. 1984 · 1984
Earlier work this paper cites.
Generalized Linear Models
Peter McCullagh and John A. Nelder. 1989 · 1989
Earlier work this paper cites.
A General Regression Neural Network
Donald F. Specht. 1991 · 1991
Earlier work this paper cites.
An introduction to kernel and nearest-neighbor nonparametric regression
Naomi S Altman. 1992 · 1992
Earlier work this paper cites.
A General Treatment of Data Redundancy in a Fuzzy Relational Data Model
G. Chen, J. Vandenbulcke, and E. E. Kerre. 1992 · 1992
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik. 1995 · 1995
Earlier work this paper cites.
Beyond Accuracy: What Data Quality Means to Data Consumers
R. Y. Wang and D. M. Strong. 1996 · 1996
Earlier work this paper cites.
OPTICS: Ordering Points to Identify the Clustering Structure. In Proceedings of the International Conference on Management of Data (SIGMOD) , Vol. 28. 49–60
Mihael Ankerst, Markus M. Breunig, Hans-Peter Kriegel, and Jörg Sander. 1999 · 1999
Earlier work this paper cites.
Comparative Accuracies of Artificial Neural Networks and Discriminant Analysis in Predicting Forest Cover Types from Cartographic Variables
J. A. Blackard and D. J. Dean. 1999 · 1999
Earlier work this paper cites.
Ridge Regression: Biased Estimation for Nonorthogonal Problems
Arthur E. Hoerl and Robert W. Kennard. 2000 · 2000
Earlier work this paper cites.
Random forests
Leo Breiman. 2001 · 2001
Earlier work this paper cites.
Greedy function approximation: A gradient boosting machine
Jerome H. Friedman. 2001 · 2001
Earlier work this paper cites.
Clustering Methods
Lior Rokach and Oded Maimon. 2005 · 2005
Earlier work this paper cites.
Data Quality: Concepts, Methods and Techniques
Carlo Batini and Monica Scannapieco. 2006 · 2006
Earlier work this paper cites.
Quality and Complexity Measures for Data Linkage and Deduplication
P. Christen and K. Goiser. 2007 · 2007
Earlier work this paper cites.
Information theoretic measures for clusterings comparison: is a correction for chance necessary?. In Proceedings of the International Conference on Machine Learning (ICML) , Vol. 382. ACM, 1073–1080
Xuan Vinh Nguyen, Julien Epps, and James Bailey. 2009 · 2009
Earlier work this paper cites.
Gaussian Mixture Models
Douglas A. Reynolds. 2009 · 2009
Earlier work this paper cites.
Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance
Xuan Vinh Nguyen, Julien Epps, and James Bailey. 2010 · 2010
Earlier work this paper cites.
Ames, Iowa: Alternative to the Boston Housing Data as an End of Semester Regression Project
Dean De Cock. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine Learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Earlier work this paper cites.
Classification in the Presence of Label Noise: A Survey
Benoît Frénay and Michel Verleysen. 2014 · 2013
Earlier work this paper cites.
An Improved k-Prototypes Clustering Algorithm for Mixed Numeric and Categorical Data
J. Ji, T. Bai, C. Zhou, C. Ma, and Z. Wang. 2013 · 2013
Cited alongside, same era.
Auto-encoder Based Data Clustering. In Iberoamerican Congress on Pattern Recognition , Vol. 8258. Springer Berlin Heidelberg, 117–124
Chunfeng Song, Feng Liu, Yongzhen Huang, Liang Wang, and Tieniu Tan. 2013 · 2013
Cited alongside, same era.
Statistical Analysis with Missing Data (second ed.)
R. J. A. Little and D. B. Rubin. 2014 · 2014
Cited alongside, same era.
A Data-driven Approach to Predict the Success of Bank Telemarketing
S. Moro, P. Cortez, and P. Rita. 2014b · 2014
Cited alongside, same era.
Efficient and Robust Automated Machine Learning. In Advances in Neural Information Processing Systems (NeurIPS) . 2962–2970
Matthias Feurer, Aaron Klein, Katharina Eggensperger, Jost Tobias Springenberg, Manuel Blum, and Frank Hutter. 2015 · 2015
Cited alongside, same era.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna M. Wallach, Hal Daumé III, and Kate Crawford. 2021 · 2021
Later among the works it cites.
There is no AI without Data
Christoph Gröger. 2021 · 2021
Later among the works it cites.
Nitin Gupta, Hima Patel, Shazia Afzal, Naveen Panwar, Ruhi Sharma Mittal, Shanmukha C. Guttula, Abhinav Jain, Lokesh Nagalapatti, Sameep Mehta, Sandeep Hans, Pranay Lohia, Aniya Aggarwal, and Diptikalyan Saha. 2021 · 2021
Later among the works it cites.
CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification Tasks. In Proceedings of the International Conference on Data Engineering (ICDE) . 13–24
Peng Li, Xi Rao, Jennifer Blase, Yue Zhang, Xu Chu, and Ce Zhang. 2021 · 2021
Later among the works it cites.
Introduction to Linear Regression Analysis
D. C. Montgomery, E. A. Peck, and G. G. Vining. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Trends in Cleaning Relational Data: Consistency and Deduplication
I. F. Ilyas and X. Chu. 2015 · 2015
Cited alongside, same era.
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Subset k-Means Approach for Handling Imbalanced-distributed Data. In Emerging ICT for Bridging the Future-Proceedings of the Annual Convention of the Computer Society of India CSI Volume 2 . Springer, 497–508
Ch. N. S. Kumar, K. N. Rao, A. Govardhan, and N. Sandhya. 2015 · 2015
Cited alongside, same era.
Applied regression: An introduction . Vol. 22
Colin Lewis-Beck and Michael Lewis-Beck. 2015 · 2015
Cited alongside, same era.
House Prices – Advanced Regression Techniques
Kaggle. 2016 · 2016
Cited alongside, same era.
PyTorch: From Research to Production
Inc. Facebook. 2017 · 2017
Cited alongside, same era.
Data quality considerations for big data and machine learning: Going beyond data cleaning and transformations
Venkat Gudivada, Amy Apon, and Junhua Ding. 2017 · 2017
Cited alongside, same era.
From Cleaning before ML to Cleaning for ML
Felix Neutatz, Binger Chen, Ziawasch Abedjan, and Eugene Wu. 2021 · 2021
Later among the works it cites.
A Data Quality-Driven View of MLOps
Cédric Renggli, Luka Rimanic, Nezihe Merve Gürel, Bojan Karlas, Wentao Wu, and Ce Zhang. 2021 · 2021
Later among the works it cites.
The Road to Software 2.0 or Data -
Chris Ré. 2021 · 2021
Later among the works it cites.
JENGA – A Framework to Study the Impact of Data Errors on the Predictions of Machine Learning Models. In Proceedings of the International Conference on Extending Database Technology (EDBT) . OpenProceedings.org, 529–534
Sebastian Schelter, Tammo Rukat, and Felix Biessmann. 2021 · 2021
Later among the works it cites.
DAG Card is the new Model Card
Jacopo Tagliabue, Ville Tuulos, Ciro Greco, and Valay Dave. 2021 · 2021
Later among the works it cites.
100,000 UK Used Car Data Set
Aditya. 2020 · 2022
Closest in time.
UCI Machine Learning Repository
D. Dua and C. Graff. 2017 · 2022
Closest in time.
Data Cleaning and AutoML: Would an Optimizer Choose to Clean?
Felix Neutatz, Binger Chen, Yazan Alkhatib, Jingwen Ye, and Ziawasch Abedjan. 2022 · 2022
Closest in time.
CUDA Toolkit
NVIDIA Corporation. visited 2022-03-11 · 2022
Closest in time.
IMDb Most Popular Films and Series
M. Ramadan. 2021 · 2022
Closest in time.
Covertype
Jock Blackard. 1998 · 2024
Closest in time.
Creditworthiness Dataset
Prof. H. Hofmann. 1994 · 2024
Closest in time.
Telco Customer Churn Dataset
IBM. 2018 · 2024
Closest in time.
Contraceptive Method Choice
T.-S. Lim. 1999 · 2024
Closest in time.
COVID-19 Dataset
Mexican Government. 2020a · 2024
Closest in time.
Información referente a casos COVID-19 en México
Mexican Government. 2020b · 2024
Closest in time.
Bank Marketing
S. Moro, P. Cortez, and P. Rita. 2014a · 2024
Closest in time.
How Do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses
Vraj Shah, Thomas Parashos, and Arun Kumar. 2024 · 2024
Closest in time.
Letter Recognition Dataset
D. J. Slate. 1991 · 2024
Closest in time.