Fetching the paper…
Reading the bibliography…
We investigate the generalization and optimization properties of shallow neural-network classifiers trained by gradient descent in the interpolating regime.
Nitanda, A., Chinot, G., and Suzuki, T. (2019) · 1905
Earlier work this paper cites.
Generalization guarantees for neural networks via harnessing the low-rank structure of the jacobian
Oymak, S., Fabian, Z., Li, M., and Soltanolkotabi, M. (2019) · 1906
Earlier work this paper cites.
For valid generalization the size of the weights is more important than the size of the network
Bartlett, P. (1996) · 1996
Earlier work this paper cites.
Almost linear vc dimension bounds for piecewise polynomial networks
Bartlett, P. L., Maiorov, V., and Meir, R. (1998) · 1998
Earlier work this paper cites.
Learning with gradient descent and weakly convex losses
Richards, D. and Rabbat, M. (2021) · 1998
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S. (2002) · 2002
Earlier work this paper cites.
Stability and generalization
Bousquet, O. and Elisseeff, A. (2002) · 2002
Earlier work this paper cites.
Norm-based capacity control in neural networks
Neyshabur, B., Tomioka, R., and Srebro, N. (2015) · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y. (2016) · 2016
Earlier work this paper cites.
Stability and generalization of learning algorithms that converge to global optima
Charles, Z. and Papailiopoulos, D. (2018) · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C. (2018) · 2018
Earlier work this paper cites.
Risk and parameter convergence of logistic regression
Ji, Z. and Telgarsky, M. (2018) · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y. (2018) · 2018
Cited alongside, same era.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D. (2018) · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N. (2018) · 2018
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S., Du, S., Hu, W., Li, Z., and Wang, R. (2019) · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F. (2019) · 2019
Cited alongside, same era.
Convergence of gradient descent on separable data
Size-independent sample complexity of neural networks
Golowich, N., Rakhlin, A., and Shamir, O. (2020) · 2020
Later among the works it cites.
Analysis of a two-layer neural network via displacement convexity
Javanmard, A., Mondelli, M., and Montanari, A. (2020) · 2020
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J. (2020) · 2020
Later among the works it cites.
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Oymak, S. and Soltanolkotabi, M. (2020) · 2020
Later among the works it cites.
When does gradient descent with logistic loss find interpolating two-layer networks?
Chatterji, N. S., Long, P. M., and Bartlett, P. L. (2021) · 2021
Later among the works it cites.
Stability & generalisation of gradient descent for shallow neural networks without the neural tangent kernel
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nacson, M. S., Lee, J., Gunasekar, S., Savarese, P. H. P., Srebro, N., and Soudry, D. (2019) · 2019
Cited alongside, same era.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Wei, C., Lee, J. D., Liu, Q., and Ma, T. (2019) · 2019
Cited alongside, same era.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Bai, Y. and Lee, J. D. (2020) · 2020
Cited alongside, same era.
Stability of stochastic gradient descent on nonsmooth convex losses
Bassily, R., Feldman, V., Guzmán, C., and Talwar, K. (2020) · 2020
Cited alongside, same era.
Generalization error bounds of gradient descent for learning over-parameterized deep relu networks
Cao, Y. and Gu, Q. (2020) · 2020
Cited alongside, same era.
How much over-parameterization is sufficient to learn deep relu networks?
Chen, Z., Cao, Y., Zou, D., and Gu, Q. (2020) · 2020
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L. and Bach, F. (2020) · 2020
Cited alongside, same era.
Richards, D. and Kuzborskij, I. (2021) · 2021
Later among the works it cites.
Gradient methods never overfit on separable data
Shamir, O. (2021) · 2021
Later among the works it cites.
Benign overfitting in two-layer convolutional neural networks
Cao, Y., Chen, Z., Belkin, M., and Gu, Q. (2022) · 2022
Later among the works it cites.
Stability and generalization analysis of gradient methods for shallow neural networks
Lei, Y., Jin, R., and Ying, Y. (2022) · 2022
Later among the works it cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Liu, C., Zhu, L., and Belkin, M. (2022) · 2022
Later among the works it cites.
Stability vs implicit bias of gradient methods on separable data and beyond
Schliserman, M. and Koren, T. (2022) · 2022
Later among the works it cites.
The sample complexity of one-hidden-layer neural networks
Vardi, G., Shamir, O., and Srebro, N. (2022) · 2022
Later among the works it cites.