Fetching the paper…
Reading the bibliography…
Random label noises (or observational noises) widely exist in practical machine learning settings.
1901
Earlier work this paper cites.
1901
Earlier work this paper cites.
S. Hochreiter, J. Schmidhuber, Flat minima, Neural Computation 9 (1) (1997) 1–42
1997
Earlier work this paper cites.
V. S. Borkar, S. K. Mitter, A strong approximation theorem for stochastic recursive algorithms, Journal of optimization theory and applications 100 (3) (1999) 499–513
1999
Earlier work this paper cites.
B. Øksendal, Stochastic differential equations, in: Stochastic differential equations, Springer, 2003, pp. 65–84
2003
Earlier work this paper cites.
2006
Earlier work this paper cites.
2006
Earlier work this paper cites.
A. Krizhevsky, G. Hinton, et al., Learning multiple layers of features from tiny images, Tech. rep. (2009)
2009
Earlier work this paper cites.
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, A. Y. Ng, Reading digits in natural images with unsupervised feature learning, in: NeurIPS Workshop on Deep Learning and Unsupervised Feature Learning, 2011
2011
Earlier work this paper cites.
A. M. Saxe, J. L. McClelland, S. Ganguli, Exact solutions to the nonlinear dynamics of learning in deep linear neural networks, in: International Conference on Learning Reporesentations (ICLR), 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. Lopez-Paz, L. Bottou, B. Schölkopf, V. Vapnik, Unifying distillation and privileged information, in: In International Conference on Learning Representations (ICLR), 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778
2016
Earlier work this paper cites.
C. Zhang, S. Bengio, M. Hardt, B. Recht, O. Vinyals, Understanding deep learning requires rethinking generalization, in: International Conference on Learning Representations, 2017
2017
Earlier work this paper cites.
E. Hoffer, I. Hubara, D. Soudry, Train longer, generalize better: closing the generalization gap in large batch training of neural networks, in: I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, R. Garnett (Eds.), Advances in Neural Information Processing Systems 30, Curran Associates, Inc., 2017, pp. 1731–1741
2017
Earlier work this paper cites.
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, P. T. P. Tang, On large-batch training for deep learning: Generalization gap and sharp minima, in: In International Conference on Learning Representations (ICLR), 2017
2017
Earlier work this paper cites.
S. Mandt, M. D. Hoffman, D. M. Blei, Stochastic gradient descent as approximate bayesian inference, The Journal of Machine Learning Research 18 (1) (2017) 4873–4907
2017
Cited alongside, same era.
Q. Li, C. Tai, E. Weinan, Stochastic modified equations and adaptive stochastic gradient algorithms, in: International Conference on Machine Learning, 2017, pp. 2101–2110
2017
Cited alongside, same era.
A. Dieuleveut, N. Flammarion, F. Bach, Harder, better, faster, stronger convergence rates for least-squares regression, The Journal of Machine Learning Research 18 (1) (2017) 3520–3570
2017
Cited alongside, same era.
L. Bottou, F. E. Curtis, J. Nocedal, Optimization methods for large-scale machine learning, Siam Review 60 (2) (2018) 223–311
2018
Cited alongside, same era.
P. Chaudhari, S. Soatto, Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks, in: International Conference on Learning Representations, 2018
H. Xiong, K. Wang, J. Bian, Z. Zhu, C.-Z. Xu, Z. Guo, J. Huan, Sphmc: Spectral hamiltonian monte carlo, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33, 2019, pp. 5516–5524
2019
Later among the works it cites.
G. Gidel, F. Bach, S. Lacoste-Julien, Implicit regularization of discrete gradient dynamics in linear neural networks, in: Advances in Neural Information Processing Systems, 2019, pp. 3202–3211
2019
Later among the works it cites.
L. Zhang, J. Song, A. Gao, J. Chen, C. Bao, K. Ma, Be your own teacher: Improve the performance of convolutional neural networks via self distillation, in: Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 3713–3722
2019
Later among the works it cites.
U. Marteau-Ferey, D. Ostrovskii, F. Bach, A. Rudi, Beyond least-squares: Fast rates for regularized empirical risk minimization through self-concordance, in: Conference on Learning Theory, 2019, pp. 2294–2340
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
B. Han, Q. Yao, X. Yu, G. Niu, M. Xu, W. Hu, I. Tsang, M. Sugiyama, Co-teaching: Robust training of deep neural networks with extremely noisy labels, in: Advances in neural information processing systems, 2018, pp. 8527–8537
2018
Cited alongside, same era.
S. L. Smith, P.-J. Kindermans, C. Ying, Q. V. Le, Don’t decay the learning rate, increase the batch size, in: International Conference on Learning Representations, 2018
2018
Cited alongside, same era.
Y. Fang, J. Xu, L. Yang, Online bootstrap confidence intervals for the stochastic gradient descent estimator, The Journal of Machine Learning Research 19 (1) (2018) 3053–3073
2018
Cited alongside, same era.
T. Li, L. Liu, A. Kyrillidis, C. Caramanis, Statistical inference using sgd., in: AAAI, 2018
2018
Cited alongside, same era.
A. Jacot, F. Gabriel, C. Hongler, Neural tangent kernel: Convergence and generalization in neural networks, in: Advances in neural information processing systems, 2018, pp. 8571–8580
2018
Cited alongside, same era.
L. Wu, C. Ma, E. Weinan, How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective, in: Advances in Neural Information Processing Systems, 2018, pp. 8279–8288
2018
Cited alongside, same era.
L. Jiang, Z. Zhou, T. Leung, L.-J. Li, L. Fei-Fei, Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels, in: International Conference on Machine Learning, 2018, pp. 2304–2313
2018
Cited alongside, same era.
Y. Wang, X. Ma, Z. Chen, Y. Luo, J. Yi, J. Bailey, Symmetric cross entropy for robust learning with noisy labels, in: Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 322–330
2019
Later among the works it cites.
J. Wu, W. Hu, H. Xiong, J. Huan, V. Braverman, Z. Zhu, On the noisy gradient descent that generalizes as sgd, in: International Conference on Machine Learning, PMLR, 2020, pp. 10367–10376
2020
Later among the works it cites.
G. Blanc, N. Gupta, G. Valiant, P. Valiant, Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process, in: Conference on learning theory, PMLR, 2020, pp. 483–513
2020
Later among the works it cites.
J. Wu, W. Hu, H. Xiong, J. Huan, V. Braverman, Z. Zhu, On the noisy gradient descent that generalizes as sgd, in: International Conference on Machine Learning (ICML), 2020
2020
Later among the works it cites.
Z. Jia, H. Su, Information-theoretic local minima characterization and regularization, in: International Conference on Machine Learning, PMLR, 2020, pp. 4773–4783
2020
Later among the works it cites.
A. Ali, E. Dobriban, R. J. Tibshirani, The implicit regularization of stochastic gradient flow for least squares, in: International Conference on Machine Learning, 2020
2020
Later among the works it cites.
Q. Xie, M.-T. Luong, E. Hovy, Q. V. Le, Self-training with noisy student improves imagenet classification, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 10687–10698
2020
Later among the works it cites.
G. Xu, Z. Liu, X. Li, C. C. Loy, Knowledge distillation meets self-supervision, in: European Conference on Computer Vision, 2020
2020
Later among the works it cites.
M. Li, M. Soltanolkotabi, S. Oymak, Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks, in: International Conference on Artificial Intelligence and Statistics, PMLR, 2020, pp. 4313–4324
2020
Later among the works it cites.
H. Xiong, X. Li, B. Yu, D. Dou, D. Wu, Z. Zhu, Implicit regularization effects of unbiased random label noises with {sgd} (2021). URL https://openreview.net/forum?id=g4szfsQUdy3
2021
Later among the works it cites.
J. Latz, Analysis of stochastic gradient descent in continuous time, Statistics and Computing 31 (4) (2021) 1–25
2021
Later among the works it cites.