Fetching the paper…
Reading the bibliography…
Recent works have cast some light on the mystery of why deep nets fit any data and generalize despite being very overparametrized.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2002
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Foundations of machine learning
Mohri, M., Rostamizadeh, A., and Talwalkar, A · 2012
Earlier work this paper cites.
Escaping from saddle points − - online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Earlier work this paper cites.
Global optimality in tensor factorization, deep learning, and beyond
Haeffele, B. D. and Vidal, R · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Neyshabur, B., Tomioka, R., and Srebro, N · 2015
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Daniely, A., Frostig, R., and Singer, Y · 2016
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J · 2016
Earlier work this paper cites.
Identity matters in deep learning
Hardt, M. and Ma, T · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kawaguchi, K · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Soudry, D. and Carmon, Y · 2016
Earlier work this paper cites.
Globally optimal gradient descent for a ConvNet with gaussian inputs
Brutzkus, A. and Globerson, A · 2017
Earlier work this paper cites.
SGD learns the conjugate kernel class of the network
Daniely, A · 2017
Earlier work this paper cites.
Dziugaite, G. K. and Roy, D. M · 2017
Earlier work this paper cites.
Size-independent sample complexity of neural networks
Golowich, N., Rakhlin, A., and Shamir, O · 2017
Earlier work this paper cites.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I · 2017
Cited alongside, same era.
PAC-Bayesian margin bounds for convolutional neural networks-technical report
Konstantinos, P., Davies, M., and Vandergheynst, P · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with ReLU activation
Li, Y. and Yuan, Y · 2017
Cited alongside, same era.
Generalization bounds of SGLD for non-convex learning: Two theoretical viewpoints
Mou, W., Wang, L., Zhai, X., and Zheng, K · 2017
Cited alongside, same era.
A PAC-Bayesian approach to spectrally-normalized margin bounds for neural networks
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N · 2017
On the power of over-parametrization in neural networks with quadratic activation
Du, S. S. and Lee, J. D · 2018
Later among the works it cites.
Deep neural networks learn non-smooth functions effectively
Imaizumi, M. and Fukumizu, K · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Later among the works it cites.
Gradient descent aligns the layers of deep linear networks
Ji, Z. and Telgarsky, M · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The loss surface of deep and wide neural networks
Nguyen, Q. and Hein, M · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
Safran, I. and Shamir, O · 2017
Cited alongside, same era.
Learning ReLUs via gradient descent
Soltanolkotabi, M · 2017
Cited alongside, same era.
Tian, Y · 2017
Cited alongside, same era.
Invariance of weight distributions in rectified mlps
Tsuchida, R., Roosta-Khorasani, F., and Gallagher, M · 2017
Cited alongside, same era.
Diverse neural network learns true target functions
Xie, B., Liang, Y., and Song, L · 2017
Cited alongside, same era.
Li, Y. and Liang, Y · 2018
Later among the works it cites.
A priori estimates of the generalization error for two-layer neural networks
Ma, C., Wu, L., et al · 2018
Later among the works it cites.
A mean field view of the landscape of two-layers neural networks
Mei, S., Montanari, A., and Nguyen, P.-M · 2018
Later among the works it cites.
Rotskoff, G. M. and Vanden-Eijnden, E · 2018
Later among the works it cites.
Mean field analysis of neural networks
Sirignano, J. and Spiliopoulos, K · 2018
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Later among the works it cites.
Neural networks with finite intrinsic dimension have no spurious valleys
Venturi, L., Bandeira, A., and Bruna, J · 2018
Later among the works it cites.
On the margin theory of feedforward neural networks
Wei, C., Lee, J. D., Liu, Q., and Ma, T · 2018
Later among the works it cites.
A critical view of global optimality in deep learning
Yun, C., Sra, S., and Jadbabaie, A · 2018
Later among the works it cites.
Learning one-hidden-layer relu networks via gradient descent
Zhang, X., Yu, Y., Wang, L., and Gu, Q · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Later among the works it cites.
The role of over-parametrization in generalization of neural networks
Neyshabur, B., Li, Z., Bhojanapalli, S., LeCun, Y., and Srebro, N · 2019
Closest in time.
Non-vacuous generalization bounds at the imagenet scale: a PAC-bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., and Orbanz, P · 2019
Closest in time.