Fetching the paper…
Reading the bibliography…
Despite their overwhelming capacity to overfit, deep neural networks trained by specific optimization algorithms tend to generalize well to unseen data.
Generalized gradients and applications
Clarke, F. H · 1975
Earlier work this paper cites.
On gradients of functions definable in o-minimal structures
Kurdyka, K · 1998
Earlier work this paper cites.
The mnist database of handwritten digits
LeCun, Y · 1998
Earlier work this paper cites.
Generalization performance of support vector machines and other pattern classifiers
Bartlett, P. and Shawe-Taylor, J · 1999
Earlier work this paper cites.
Data Mining: Practical Machine Learning Tools and Techniques, (Morgan Kaufmann Series in Data Management Systems)
Witten, I. H. and Frank, E · 2005
Earlier work this paper cites.
Real analysis: measure theory, integration, and Hilbert spaces
Stein, E. M. and Shakarchi, R · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini–batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
New types of deep neural network learning for speech recognition and related applications: An overview
Deng, L., Hinton, G., and Kingsbury, B · 2013
Earlier work this paper cites.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Path-sgd: Path-normalized optimization in deep neural networks
Neyshabur, B., Salakhutdinov, R. R., and Srebro, N · 2015
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Ruder, S · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M · 2017
Cited alongside, same era.
Improving generalization performance by switching from adam to sgd
Keskar, N. S. and Socher, R · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2018
Cited alongside, same era.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
Chen, J., Zhou, D., Tang, Y., Yang, Z., Cao, Y., and Gu, Q · 2018
Cited alongside, same era.
When will gradient methods converge to max-margin classifier under relu models?
Xu, T., Zhou, Y., Ji, K., and Liang, Y · 2018
Later among the works it cites.
Recent trends in deep learning based natural language processing
Young, T., Hazarika, D., Poria, S., and Cambria, E · 2018
Later among the works it cites.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., and Wang, R · 2019
Later among the works it cites.
The implicit bias of gradient descent on nonseparable data
Ji, Z. and Telgarsky, M · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S · 2019
Later among the works it cites.
Implicit bias of gradient descent based adversarial training on separable data
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient descent aligns the layers of deep linear networks
Ji, Z. and Telgarsky, M · 2018
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2018
Cited alongside, same era.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L., Xiong, Y., Liu, Y., and Sun, X · 2018
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Cited alongside, same era.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Neyshabur, B., Bhojanapalli, S., and Srebro, N · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
Reddi, S., Zaheer, M., Sachan, D., Kale, S., and Kumar, S · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Cited alongside, same era.
Li, Y., Fang, E. X., Xu, H., and Zhao, T · 2019
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J · 2019
Later among the works it cites.
The implicit bias of adagrad on separable data
Qian, Q. and Qian, X · 2019
Later among the works it cites.
Stochastic subgradient method converges on tame functions
Davis, D., Drusvyatskiy, D., Kakade, S., and Lee, J. D · 2020
Closest in time.
Directional convergence and alignment in deep learning
Ji, Z. and Telgarsky, M · 2020
Closest in time.
Towards theoretically understanding why sgd generalizes better than adam in deep learning
Zhou, P., Feng, J., Ma, C., Xiong, C., Hoi, S. C. H., et al · 2020
Closest in time.
Adabelief optimizer: Adapting stepsizes by the belief in observed gradients
Zhuang, J., Tang, T., Ding, Y., Tatikonda, S. C., Dvornek, N., Papademetris, X., and Duncan, J · 2020
Closest in time.