Fetching the paper…
Reading the bibliography…
The optimization of multilayer neural networks typically leads to a solution with zero training error, yet the landscape can exhibit spurious local minima and the minima can be disconnected.
Mean field limit of the learning dynamics of multilayer neural networks
Nguyen, P.-M · 1902
Earlier work this paper cites.
Mean field analysis of deep neural networks
Sirignano, J. and Spiliopoulos, K · 1903
Earlier work this paper cites.
Training a 3-node neural network is NP-complete
Blum, A. and Rivest, R. L · 1989
Earlier work this paper cites.
Exponentially many local minima for single neurons
Auer, P., Herbster, M., and Warmuth, M. K · 1996
Earlier work this paper cites.
A rigorous framework for the mean field limit of multilayer neural networks
Nguyen, P.-M. and Pham, H. T · 2001
Earlier work this paper cites.
A note on the global convergence of multilayer neural networks in the mean field regime
Nguyen, P.-M. and Pham, H. T · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B · 2008
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., and Shamir, O · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J., Vinyals, O., and Saxe, A. M · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Earlier work this paper cites.
The non-convex burer-monteiro approach works on smooth semidefinite programs
Boumal, N., Voroninski, V., and Bandeira, A · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Ge, R., Lee, J. D., and Ma, T · 2016
Earlier work this paper cites.
Deep learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Earlier work this paper cites.
Dynamic network surgery for efficient DNNs
Guo, Y., Yao, A., and Chen, Y · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kawaguchi, K · 2016
Cited alongside, same era.
On the quality of the initial basin in overspecified neural networks
Safran, I. and Shamir, O · 2016
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J · 2017
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Ge, R., Jin, C., and Zheng, Y · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Cited alongside, same era.
Variational dropout sparsifies deep neural networks
Molchanov, D., Ashukha, A., and Vetrov, D · 2017
Cited alongside, same era.
Mean field analysis of neural networks
Sirignano, J. and Spiliopoulos, K · 2018
Later among the works it cites.
On the margin theory of feedforward neural networks
Wei, C., Lee, J. D., Liu, Q., and Ma, T · 2018
Later among the works it cites.
A critical view of global optimality in deep learning
Yun, C., Sra, S., and Jadbabaie, A · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Later among the works it cites.
A mean-field limit for certain deep neural networks
Araújo, D., Oliveira, R. I., and Yukimura, D · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The loss surface of deep and wide neural networks
Nguyen, Q. and Hein, M · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
Pennington, J. and Bahri, Y · 2017
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L. and Bach, F · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F · 2018
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S. S., Lee, J. D., Li, H., Wang, L., and Zhai, X · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Cited alongside, same era.
Closest in time.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2019
Closest in time.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2019
Closest in time.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 2019
Closest in time.
Analysis of a two-layer neural network via displacement convexity
Javanmard, A., Mondelli, M., and Montanari, A · 2019
Closest in time.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Kuditipudi, R., Wang, X., Lee, H., Zhang, Y., Li, Z., Hu, W., Arora, S., and Ge, R · 2019
Closest in time.
Rethinking the value of network pruning
Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T · 2019
Closest in time.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Mei, S., Misiakiewicz, T., and Montanari, A · 2019
Closest in time.
On the loss landscape of a class of deep neural networks with no bad local valleys
Nguyen, Q., Mukkamala, M. C., and Hein, M · 2019
Closest in time.
Spurious valleys in two-layer neural network optimization landscapes
Venturi, L., Bandeira, A. S., and Bruna, J · 2019
Closest in time.
Mean-field analysis of two-layer neural networks: Non-asymptotic rates and generalization bounds
Chen, Z., Cao, Y., Gu, Q., and Zhang, T · 2020
Closest in time.