Fetching the paper…
Reading the bibliography…
With an eye toward understanding complexity control in deep learning, we study how infinitesimal regularization or gradient descent optimization lead to margin maximizing solutions in both homogeneous and non-homogeneous models, extending previous work that focused on infinitesimal regularization only in homogeneous models.
Numerical optimization
Nocedal, J. and Wright, S · 2006
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B. E., Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Earlier work this paper cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Earlier work this paper cites.
Geometry of optimization and implicit regularization in deep learning
Neyshabur, B., Tomioka, R., Salakhutdinov, R., and Srebro, N · 2017
Earlier work this paper cites.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Cited alongside, same era.
Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
Wu, L., Zhu, Z., and E, W · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Cited alongside, same era.
A continuous-time view of early stopping for least squares regression
Ali, A., Kolter, J. Z., and Tibshirani, R. J · 2018
Cited alongside, same era.
Connecting optimization and regularization paths
Suggala, A., Prasad, A., and Ravikumar, P. K · 2018
Later among the works it cites.
On the margin theory of feedforward neural networks
Wei, C., Lee, J. D., Liu, Q., and Ma, T · 2018
Later among the works it cites.
Gradient descent aligns the layers of deep linear networks
Ji, Z. and Telgarsky, M · 2019
Closest in time.
Cross-entropy loss leads to poor margins, 2019
Nar, K., Ocal, O., Sastry, S. S., and Ramchandran, K · 2019
Closest in time.
When will gradient methods converge to max-margin classifier under reLU models?, 2019
Xu, T., Zhou, Y., Ji, K., and Liang, Y · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Byrd, J. and Lipton, Z. C · 2018
Cited alongside, same era.
Risk and parameter convergence of logistic regression
Ji, Z. and Telgarsky, M · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J., Soudry, D., and Srebro, N
Cited in the paper.
Characterizing implicit bias in terms of optimization geometry
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N
Cited in the paper.
Convergence of gradient descent on separable data
Nacson, M. S., Lee, J., Gunasekar, S., Savarese, P. H., Srebro, N., and Soudry, D
Cited in the paper.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
Nacson, M. S., Srebro, N., and Soudry, D
Cited in the paper.
Path-SGD: Path-normalized optimization in deep neural networks
Neyshabur, B., Salakhutdinov, R. R., and Srebro, N
Cited in the paper.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N
Cited in the paper.