Fetching the paper…
Reading the bibliography…
We study how permutation symmetries in overparameterized multi-layer neural networks generate `symmetry-induced' critical points.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S. S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J · 1902
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
Fukumizu, K. and Amari, S.-i · 2000
Earlier work this paper cites.
The calculus of finite differences
Milne-Thomson, L. M · 2000
Earlier work this paper cites.
Statistics of critical points of gaussian fields on large-dimensional spaces
Bray, A. J. and Dean, D. S · 2007
Earlier work this paper cites.
Combinatorics: the Rota way
Kung, J. P., Rota, G.-C., and Yan, C. H · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Random matrices and complexity of spin glasses
Auffinger, A., Arous, G. B., and Černỳ, J · 2013
Earlier work this paper cites.
The symmetric group: representations, combinatorial algorithms, and symmetric functions , volume 203
Sagan, B. E · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Explorations on high dimensional landscapes
Sagun, L., Guney, V. U., Arous, G. B., and LeCun, Y · 2014
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J · 2016
Earlier work this paper cites.
Gradient descent can take exponential time to escape saddle points
Du, S. S., Jin, C., Lee, J. D., Jordan, M. I., Singh, A., and Poczos, B · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2017
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L. and Bach, F · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. A · 2018
Large scale structure of neural network loss landscapes
Fort, S. and Jastrzebski, S · 2019
Later among the works it cites.
Semi-flat minima and saddle points by embedding neural networks to overparameterization
Fukumizu, K., Yamaguchi, S., Mototake, Y.-i., and Tanaka, M · 2019
Later among the works it cites.
Jamming transition as a paradigm to understand the loss landscape of deep neural networks
Geiger, M., Spigler, S., d’Ascoli, S., Sagun, L., Baity-Jesi, M., Biroli, G., and Wyart, M · 2019
Later among the works it cites.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Kuditipudi, R., Wang, X., Lee, H., Zhang, Y., Li, Z., Hu, W., Ge, R., and Arora, S · 2019
Later among the works it cites.
On connected sublevel sets in deep learning
Nguyen, Q · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Arora, S., Du, S. S., Hu, W., Li, Z., and Wang, R · 2019
Cited alongside, same era.
Brea, J., Simsek, B., Illing, B., and Gerstner, W · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S., Lee, J., Li, H., Wang, L., and Zhai, X · 2019
Cited alongside, same era.
The critical locus of overparameterized neural networks
Cooper, Y · 2020
Later among the works it cites.
Fort, S., Dziugaite, G. K., Paul, M., Kharaghani, S., Roy, D. M., and Ganguli, S · 2020
Later among the works it cites.
Linear mode connectivity and the lottery ticket hypothesis
Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M · 2020
Later among the works it cites.
Neural mechanics: Symmetry and broken conservation laws in deep learning dynamics
Kunin, D., Sagastuy-Brena, J., Ganguli, S., Yamins, D. L., and Tanaka, H · 2020
Later among the works it cites.
Finite versus infinite neural networks: an empirical study
Lee, J., Schoenholz, S. S., Pennington, J., Adlam, B., Xiao, L., Novak, R., and Sohl-Dickstein, J · 2020
Later among the works it cites.
Genni: Visualising the geometry of equivalences for neural network identifiability
Lengyel, D., Petangoda, J., Falk, I., Highnam, K., Lazarou, M., Kolbeinsson, A., Deisenroth, M. P., and Jennings, N. R · 2020
Later among the works it cites.
Noether: The more things change, the more stay the same
Głuch, G. and Urbanke, R · 2021
Closest in time.