Fetching the paper…
Reading the bibliography…
Recent theoretical results show that gradient descent on deep neural networks under exponential loss functions locally maximizes classification margin, which is equivalent to minimizing the norm of the weight matrices under margin constraints.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J · 1906
Earlier work this paper cites.
Statistical learning theory
Vapnik, V. N · 1998
Earlier work this paper cites.
Almost-everywhere algorithmic stability and generalization error
Kutin, S. and Niyogi, P · 2002
Earlier work this paper cites.
(not) bounding the true error
Langford, J. and Caruana, R · 2002
Earlier work this paper cites.
Introduction to statistical learning theory
Bousquet, O., Boucheron, S., and Lugosi, G · 2003
Earlier work this paper cites.
Learning theory: stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization
Mukherjee, S., Niyogi, P., Poggio, T., and Rifkin, R · 2006
Earlier work this paper cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Bartlett, P., Foster, D. J., and Telgarsky, M · 2017
Earlier work this paper cites.
The Implicit Bias of Gradient Descent on Separable Data
Soudry, D., Hoffer, E., and Srebro, N · 2017
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2018
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Arora, S., Ge, R., Neyshabur, B., and Zhang, Y · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S. S., Lee, J. D., Li, H., Wang, L., and Zhai, X · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Cited alongside, same era.
A surprising linear relationship predicts test performance in deep networks, 2018
Liao, Q., Miranda, B., Banburski, A., Hidary, J., and Poggio, T · 2018
Theory of deep learning III: Dynamics and generalization in deep networks
Banburski, A., Liao, Q., Miranda, B., Poggio, T., Rosasco, L., Liang, B., and Hidary, J · 2019
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2019
Later among the works it cites.
Generalization in deep network classifiers trained with the square loss
Poggio, T. and Liao, Q · 2019
Later among the works it cites.
Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models
Shpigel Nacson, M., Gunasekar, S., Lee, J. D., Srebro, N., and Soudry, D · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the importance of single directions for generalization
Morcos, A. S., Barrett, D. G., Rabinowitz, N. C., and Botvinick, M · 2018
Cited alongside, same era.
Wang, T., Zhu, J., Torralba, A., and Efros, A. A · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Cited alongside, same era.
Arora, S., Du, S. S., Hu, W., Li, Z., and Wang, R · 2019
Cited alongside, same era.
Theoretical issues in deep networks
Poggio, T., Banburski, A., and Liao, Q
Cited in the paper.
Complexity control by gradient descent in deep networks
Poggio, T., Liao, Q., and Banburski, A
Cited in the paper.
Prevalence of neural collapse during the terminal phase of deep learning training
Papyan, V., Han, X. Y., and Donoho, D. L · 2020
Later among the works it cites.
Loss landscape: Sgd can have a better view than gd
Poggio, T. and Cooper, Y · 2020
Later among the works it cites.
Implicit dynamic regularization in deep networks
Poggio, T. A. and Liao, Q · 2020
Later among the works it cites.
Deep learning on a data diet: Finding important examples early in training, 2021
Paul, M., Ganguli, S., and Dziugaite, G. K · 2021
Closest in time.