Fetching the paper…
Reading the bibliography…
Training modern neural networks is an inherently noisy process that can lead to high \emph{prediction churn} -- disagreements between re-trainings of the same model due to factors such as randomization in the parameter initialization and mini-batches -- even when the trained models all attain similar accuracies.
Discriminatory analysis-nonparametric discrimination: consistency properties
Fix, E. and Hodges Jr, J. L · 1951
Earlier work this paper cites.
Rates of convergence for nearest neighbor procedures
Cover, T. M · 1968
Earlier work this paper cites.
Consistent nonparametric regression
Stone, C. J · 1977
Earlier work this paper cites.
On the strong universal consistency of nearest neighbor regression function estimates
Devroye, L., Gyorfi, L., Krzyzak, A., Lugosi, G., et al · 1994
Earlier work this paper cites.
On nonparametric estimation of density level sets
Tsybakov, A. B. et al · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Adaptive Hausdorff estimation of density level sets
Singh, A., Scott, C., Nowak, R., et al · 2009
Earlier work this paper cites.
Rates of convergence for the cluster tree
Chaudhuri, K. and Dasgupta, S · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
A comparative analysis of offline and online evaluations and discussion of research paper recommender system evaluation
Beel, J., Genzmehr, M., Langer, S., Nürnberger, A., and Gipp, B · 2013
Earlier work this paper cites.
Improving the sensitivity of online controlled experiments by utilizing pre-experiment data
Deng, A., Xu, Y., Kohavi, R., and Walker, T · 2013
Earlier work this paper cites.
Rates of convergence for nearest neighbor classification
Chaudhuri, K. and Dasgupta, S · 2014
Earlier work this paper cites.
Objective bayesian two sample hypothesis testing for online controlled experiments
Deng, A · 2015
Earlier work this paper cites.
Online batch selection for faster training of neural networks
Loshchilov, I. and Hutter, F · 2015
Earlier work this paper cites.
Ad recommendation systems for life-time value optimization
Theocharous, G., Thomas, P. S., and Ghavamzadeh, M · 2015
Earlier work this paper cites.
Estimating numerical error in neural network simulations on graphics processing units
Turner, J. P. and Nowotny, T · 2015
Cited alongside, same era.
Data-driven metric development for online controlled experiments: Seven lessons learned
Deng, A. and Shi, X · 2016
Cited alongside, same era.
Measuring metrics
Dmitriev, P. and Wu, X · 2016
Cited alongside, same era.
Launch and iterate: Reducing prediction churn
Fard, M. M., Cormier, Q., Canini, K., and Gupta, M · 2016
Cited alongside, same era.
Satisfying real-world goals with dataset constraints
Goh, G., Cotter, A., Gupta, M., and Friedlander, M. P · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Cited alongside, same era.
Large-scale celebfaces attributes (celeba) dataset
Liu, Z., Luo, P., Wang, X., and Tang, X · 2018
Later among the works it cites.
Deep k-nearest neighbors: Towards confident, interpretable and robust deep learning
Papernot, N. and McDaniel, P · 2018
Later among the works it cites.
How does batch normalization help optimization?
Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A · 2018
Later among the works it cites.
Collaborative learning for deep neural networks
Song, G. and Chai, W · 2018
Later among the works it cites.
Deep mutual learning
Zhang, Y., Xiang, T., Hospedales, T. M., and Lu, H · 2018
Later among the works it cites.
Robust bi-tempered logistic loss based on bregman divergences
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving the robustness of deep neural networks via stability training
Zheng, S., Song, Y., Leung, T., and Goodfellow, I · 2016
Cited alongside, same era.
UCI machine learning repository, 2017
Dua, D. and Graff, C · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Cited alongside, same era.
Decoupling” when to update” from” how to update”
Malach, E. and Shalev-Shwartz, S · 2017
Cited alongside, same era.
Randomness in neural networks: an overview
Scardapane, S. and Wang, D · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2017
Cited alongside, same era.
Amid, E., Warmuth, M. K., Anil, R., and Koren, T · 2019
Later among the works it cites.
Optimization with non-differentiable constraints with applications to fairness, recall, churn, and other goals
Cotter, A., Jiang, H., Gupta, M. R., Wang, S., Narayan, T., You, S., and Sridharan, K · 2019
Later among the works it cites.
Deep ensembles: A loss landscape perspective
Fort, S., Hu, H., and Lakshminarayanan, B · 2019
Later among the works it cites.
Non-asymptotic uniform rates of consistency for k-nn regression
Jiang, H · 2019
Later among the works it cites.
When does label smoothing help?
Müller, R., Kornblith, S., and Hinton, G. E · 2019
Later among the works it cites.
Fast rates for a kNN classifier robust to unknown asymmetric label noise
Reeve, H. W. and Kaban, A · 2019
Later among the works it cites.
A survey on image data augmentation for deep learning
Shorten, C. and Khoshgoftaar, T. M · 2019
Later among the works it cites.
Combating label noise in deep learning using abstention
Thulasidasan, S., Bhattacharya, T., Bilmes, J., Chennupati, G., and Mohd-Yusof, J · 2019
Later among the works it cites.
Knowledge distillation by on-the-fly native ensemble
Zhu, X., Gong, S., et al · 2019
Later among the works it cites.
Deep k-nn for noisy labels
Bahri, D., Jiang, H., and Gupta, M · 2020
Later among the works it cites.