Fetching the paper…
Reading the bibliography…
In our experience of working with domain experts who are using today's AutoML systems, a common problem we encountered is what we call "unrealistic expectations" -- when users are facing a very challenging task with a noisy data acquisition process, while being expected to achieve startlingly high accuracy with machine learning (ML).
T. M. Cover and P. A. Hart, “Nearest neighbor pattern classification,” IEEE Transactions on Information Theory , vol. 13, no. 1, pp. 21–27, 1967
1967
Earlier work this paper cites.
K. Fukunaga and D. Kessell, “Nonparametric Bayes error estimation using unclassified samples,” IEEE Transactions on Information Theory , vol. 19, no. 4, pp. 434–440, 1973
1973
Earlier work this paper cites.
K. Fukunaga and L. Hostetler, “k-nearest-neighbor Bayes-risk estimation,” IEEE Transactions on Information Theory , vol. 21, no. 3, pp. 285–293, 1975
1975
Earlier work this paper cites.
P. A. Devijver, “A multiclass, k-NN approach to Bayes risk estimation,” Pattern recognition letters , vol. 3, no. 1, pp. 1–6, 1985
1985
Earlier work this paper cites.
K. Fukunaga and D. M. Hummels, “Bayes error estimation using parzen and k-NN procedures,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 9, no. 5, pp. 634–643, May 1987
1987
Earlier work this paper cites.
R. R. Snapp, D. Psaltis, and S. S. Venkatesh, “Asymptotic slowing down of the nearest-neighbor classifier,” in Advances in Neural Information Processing Systems , 1991, pp. 932–938
1991
Earlier work this paper cites.
L. J. Buturovic and M. Z. Markovic, “Improving k-nearest neighbor bayes error estimates,” in 11th IAPR International Conference on Pattern Recognition. Vol. II. Conference B: Pattern Recognition Methodology and Systems , vol. 1. IEEE Computer Society, 1992, pp. 470–471
1992
Earlier work this paper cites.
R. Y. Wang and D. M. Strong, “Beyond accuracy: What data quality means to data consumers,” Journal of management information systems , vol. 12, no. 4, pp. 5–33, 1996
1996
Earlier work this paper cites.
R. R. Snapp and T. Xu, “Estimating the Bayes risk from sample data,” in Advances in Neural Information Processing Systems , 1996, pp. 232–238
1996
Earlier work this paper cites.
D. M. Strong, Y. W. Lee, and R. Y. Wang, “Data quality in context,” Communications of the ACM , vol. 40, no. 5, 1997
1997
Earlier work this paper cites.
M. Scannapieco and T. Catarci, “Data quality under a computer science perspective,” Archivi & Computer , vol. 2, 2002
2002
Earlier work this paper cites.
G. Cong, W. Fan, F. Geerts, X. Jia, and S. Ma, “Improving data quality: Consistency and accuracy.” in VLDB , vol. 7, 2007, pp. 315–326
2007
Earlier work this paper cites.
T. Pham-Gia, N. Turkkan, and A. Bekker, “Bounds for the Bayes error in classification: A Bayesian approach using discriminant analysis,” Statistical Methods & Applications , vol. 16, no. 1, pp. 7–26, Jun. 2007
2007
Earlier work this paper cites.
H. Van Vliet, H. Van Vliet, and J. Van Vliet, Software engineering: principles and practice . John Wiley & Sons, 2008, vol. 13
2008
Earlier work this paper cites.
F. Chiang and R. J. Miller, “Discovering data quality rules,” Proceedings of the VLDB Endowment , vol. 1, no. 1, pp. 1166–1177, 2008
2008
Earlier work this paper cites.
C. Batini, C. Cappiello, C. Francalanci, and A. Maurino, “Methodologies for data quality assessment and improvement,” ACM computing surveys , vol. 41, no. 3, 2009
2009
Earlier work this paper cites.
S. Sadiq, N. K. Yeganeh, and M. Indulska, “20 years of data quality research: themes, trends and synergies,” in Proceedings of the Twenty-Second Australasian Database Conference-Volume 115 , 2011, pp. 153–162
2011
Earlier work this paper cites.
D. Compton, T. P. Love, and J. Sell, “Developing and assessing intercoder reliability in studies of group interaction,” Sociological Methodology , vol. 42, no. 1, pp. 348–364, 2012
2012
Earlier work this paper cites.
J. Bergstra, D. Yamins, and D. D. Cox, “Hyperopt: A python library for optimizing the hyperparameters of machine learning algorithms,” in Proceedings of the 12th Python in Science Conference , vol. 13. Citeseer, 2013, p. 20
2013
Earlier work this paper cites.
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B. Su, “Scaling distributed machine learning with the parameter server,” in OSDI , 2014, pp. 583–598
2014
Earlier work this paper cites.
W. Fan, “Data quality: From theory to practice,” Acm Sigmod Record , vol. 44, no. 3, pp. 7–18, 2015
2015
Earlier work this paper cites.
U. Gadiraju, R. Kawase, S. Dietze, and G. Demartini, “Understanding malicious behavior in crowdsourcing platforms: The case of online surveys,” in Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems , 2015, pp. 1631–1640
2015
Earlier work this paper cites.
X. Meng, J. K. Bradley, B. Yavuz, E. R. Sparks, S. Venkataraman, D. Liu, J. Freeman, D. B. Tsai, M. Amde, S. Owen, D. Xin, R. Xin, M. J. Franklin, R. Zadeh, M. Zaharia, and A. Talwalkar, “MLlib: Machine learning in Apache Spark,” Journal of Machine Learning Research , vol. 17, pp. 34:1–34:7, 2016
2016
Earlier work this paper cites.
M. Vartak, H. Subramanyam, W.-E. Lee, S. Viswanathan, S. Husnoo, S. Madden, and M. Zaharia, “Modeldb: a system for machine learning model management,” in Proceedings of the Workshop on Human-In-the-Loop Data Analytics , 2016, pp. 1–3
2016
Cited alongside, same era.
Z. Abedjan, X. Chu, D. Deng, R. C. Fernandez, I. F. Ilyas, M. Ouzzani, P. Papotti, M. Stonebraker, and N. Tang, “Detecting data errors: Where are we and what needs to be done?” Proceedings of the VLDB Endowment , vol. 9, no. 12, pp. 993–1004, 2016
2016
Cited alongside, same era.
S. Krishnan, J. Wang, E. Wu, M. J. Franklin, and K. Goldberg, “ActiveClean: Interactive Data Cleaning for Statistical Modeling,” Proceedings of the VLDB Endowment , vol. 9, no. 12, 2016
2016
Cited alongside, same era.
V. Berisha, A. Wisler, A. O. Hero, and A. Spanias, “Empirically estimable classification bounds based on a nonparametric divergence measure,” IEEE Transactions on Signal Processing , vol. 64, no. 3, pp. 580–591, 2016
2016
2019
Later among the works it cites.
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” in Advances in Neural Information Processing Systems , 2019, pp. 5754–5764
2019
Later among the works it cites.
M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in International Conference on Machine Learning . PMLR, 2019, pp. 6105–6114
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
K. Jamieson and A. Talwalkar, “Non-stochastic best arm identification and hyperparameter optimization,” in Artificial Intelligence and Statistics , 2016, pp. 240–248
2016
Cited alongside, same era.
D. Baylor, E. Breck, H.-T. Cheng, N. Fiedel, C. Y. Foo, Z. Haque, S. Haykal, M. Ispir, V. Jain, L. Koc et al. , “Tfx: A tensorflow-based production-scale machine learning platform,” in Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , 2017, pp. 1387–1395
2017
Cited alongside, same era.
L. Li, K. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar, “Hyperband: A novel bandit-based approach to hyperparameter optimization,” The Journal of Machine Learning Research , vol. 18, no. 1, pp. 6765–6816, 2017
2017
Cited alongside, same era.
M. Zaharia, A. Chen, A. Davidson, A. Ghodsi, S. A. Hong, A. Konwinski, S. Murching, T. Nykodym, P. Ogilvie, M. Parkhe et al. , “Accelerating the machine learning lifecycle with mlflow.” IEEE Data Eng. Bull. , vol. 41, no. 4, pp. 39–45, 2018
2018
Cited alongside, same era.
T. Kraska, “Northstar: An interactive data science system,” PVLDB , vol. 11, no. 12, pp. 2150–2164, 2018
2018
Cited alongside, same era.
S. Schelter, D. Lange, P. Schmidt, M. Celikel, F. Biessmann, and A. Grafberger, “Automating large-scale data quality verification,” Proceedings of the VLDB Endowment , vol. 11, no. 12, pp. 1781–1794, 2018
2018
Cited alongside, same era.
A. Checco, J. Bates, and G. Demartini, “All that glitters is gold—an attack scheme on gold questions in crowdsourcing,” in Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , vol. 6, no. 1, 2018
2018
Cited alongside, same era.
Z. Wu, A. A. Efros, and S. X. Yu, “Improving generalization via scalable neighborhood component analysis,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 685–701
2018
Cited alongside, same era.
S. Nakandala, Y. Zhang, and A. Kumar, “Cerebro: a data system for optimized deep learning model selection,” Proceedings of the VLDB Endowment , vol. 13, no. 12, pp. 2159–2173, 2020
2020
Closest in time.
2020
Closest in time.
W. Wu, L. Flokas, E. Wu, and J. Wang, “Complaint-driven training data debugging for query 2.0,” in Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data , 2020, pp. 1317–1334
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
S. Y. Sekeh, B. L. Oselio, and A. O. Hero, “Learning to bound the multi-class Bayes error,” IEEE Transactions on Signal Processing , 2020
2020
Closest in time.
L. Rimanic, C. Renggli, B. Li, and C. Zhang, “On convergence of nearest neighbor classifiers over feature transformations,” Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Closest in time.
2020
Closest in time.
J. S. Rosenfeld, A. Rosenfeld, Y. Belinkov, and N. Shavit, “A constructive prediction of the generalization error across scales,” in International Conference on Learning Representations , 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur, “Sharpness-aware minimization for efficiently improving generalization,” in International Conference on Learning Representations , 2020
2020
Closest in time.
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He, “Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,” in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2020, pp. 3505–3506
2020
Closest in time.
C. Renggli, L. Rimanic, N. M. Gürel, B. Karlaš, W. Wu, and C. Zhang, “A data quality-driven view of mlops,” 2021
2021
Closest in time.
C. Renggli, L. Rimanic, N. Hollenstein, and C. Zhang, “Evaluating bayes error estimators on read-world datasets with feebee,” Advances in Neural Information Processing Systems (Datasets and Benchmarks) , vol. 34, 2021
2021
Closest in time.
T. Hashimoto, “Model performance scaling with multiple data sources,” in International Conference on Machine Learning . PMLR, 2021, pp. 4107–4116
2021
Closest in time.
A. Byerly, T. Kalganova, and I. Dear, “No routing needed between capsules,” Neurocomputing , 2021
2021
Closest in time.
2021
Closest in time.
J. Wei, Z. Zhu, H. Cheng, T. Liu, G. Niu, and Y. Liu, “Learning with noisy labels revisited: A study using real-world human annotations,” in International Conference on Learning Representations , 2022
2022
Closest in time.