Fetching the paper…
Reading the bibliography…
Over-parameterization and adaptive methods have played a crucial role in the success of deep learning in the last decade.
Global convergence of adaptive gradient methods for an over-parameterized neural network
Wu, X., Du, S. S., and Ward, R. (2019) · 1902
Earlier work this paper cites.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J. (2019) · 1903
Earlier work this paper cites.
On the convergence of adam and beyond
Reddi, S. J., Kale, S., and Kumar, S. (2019) · 1904
Earlier work this paper cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S. and Montanari, A. (2019) · 1908
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
The elements of statistical learning
Friedman, J., Hastie, T., and Tibshirani, R. (2001) · 2001
Earlier work this paper cites.
Neural networks: tricks of the trade
Orr, G. and Müller, K.-R. (2003) · 2003
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C. M. (2006) · 2006
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L. (2010) · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Earlier work this paper cites.
Towards optimal one pass large scale learning with averaged stochastic gradient descent
Xu, W. (2011) · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Bengio, Y. (2012) · 2012
Earlier work this paper cites.
Stochastic gradient descent tricks
Bottou, L. (2012) · 2012
Earlier work this paper cites.
Lecture 6.5-RMSPro: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G. (2012) · 2012
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
Zeiler, M. (2012) · 2012
Cited alongside, same era.
No more pesky learning rates
Schaul, T., Zhang, S., and LeCun, Y. (2013) · 2013
Cited alongside, same era.
An empirical study of learning rates in deep neural networks for speech recognition
Senior, A., Heigold, G., and Yang, K. (2013) · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013) · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J. (2014) · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Variants of RMSProp and AdaGrad with logarithmic regret bounds
Mukkamala, M. C. and Hein, M. (2017) · 2017
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P.-J., and Le, Q. V. (2017) · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Wilson, A., Roelofs, R., Stern, M., Srebro, N., and Recht, B. (2017) · 2017
Later among the works it cites.
Zhang, J. and Mitliagkas, I. (2017) · 2017
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z. (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shalev-Shwartz, S. and Ben-David, S. (2014) · 2014
Cited alongside, same era.
Enhanced image classification with a fast-learning shallow convolutional neural network
McDonnell, M. D. and Vladusich, T. (2015) · 2015
Cited alongside, same era.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015) · 2015
Cited alongside, same era.
Incorporating Nesterov momentum into Adam
Dozat, T. (2016) · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Ruder, S. (2016) · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Telgarsky, M. (2016) · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2016) · 2016
Cited alongside, same era.
Later among the works it cites.
Minnorm training: an algorithm for training over-parameterized deep neural networks
Bansal, Y., Advani, M., Cox, D., and Saxe, A. (2018) · 2018
Later among the works it cites.
To understand deep learning we need to understand kernel learning
Belkin, M., Ma, S., and Mandal, S. (2018) · 2018
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
Du, S. S. and Lee, J. D. (2018) · 2018
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A. (2018) · 2018
Later among the works it cites.
Characterizing implicit bias in terms of optimization geometry
Gunasekar, S., Lee, J., Soudry, D., and Srebro, N. (2018) · 2018
Later among the works it cites.
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. (2018) · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N. (2018) · 2018
Later among the works it cites.
On the convergence proof of amsgrad and a new version
Tran, P. T. et al. (2019) · 2019
Later among the works it cites.
Benign overfitting in linear regression
Bartlett, P. L., Long, P. M., Lugosi, G., and Tsigler, A. (2020) · 2020
Closest in time.