Fetching the paper…
Reading the bibliography…
This work finds the analytical expression of the global minima of a deep linear network with weight decay and stochastic neurons, a fundamental model for understanding the landscape of neural networks.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J. (2019) · 1903
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Efficient capital markets: A review of theory and empirical work
Fama, E. F. (1970) · 1970
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K. (1989) · 1989
Earlier work this paper cites.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J. A. (1992) · 1992
Earlier work this paper cites.
Bayesian methods for adaptive models
Mackay, D. J. C. (1992) · 1992
Earlier work this paper cites.
Piecewise linear activations substantially shape the loss surfaces of neural networks
He, F., Wang, B., and Tao, D. (2020) · 2003
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M. (2013) · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S. (2013) · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z. (2016) · 2016
Earlier work this paper cites.
Identity matters in deep learning
Hardt, M. and Ma, T. (2016) · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kawaguchi, K. (2016) · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2017) · 2017
Cited alongside, same era.
Depth creates no bad local minima
Lu, H. and Kawaguchi, K. (2017) · 2017
Cited alongside, same era.
Searching for activation functions
Ramachandran, P., Zoph, B., and Le, Q. V. (2017) · 2017
Cited alongside, same era.
How regularization affects the critical points in linear networks
Taghvaei, A., Kim, J. W., and Mehta, P. (2017) · 2017
Cited alongside, same era.
Fixing a broken ELBO
Alemi, A., Poole, B., Fischer, I., Dillon, J., Saurous, R. A., and Murphy, K. (2018) · 2018
Cited alongside, same era.
Dropout as a low-rank regularizer for matrix factorization
Cavazza, J., Morerio, P., Haeffele, B., Lane, C., Murino, V., and Vidal, R. (2018) · 2018
Pruning neural networks without any data by iteratively conserving synaptic flow
Tanaka, H., Kunin, D., Yamins, D. L., and Ganguli, S. (2020) · 2020
Later among the works it cites.
Spurious local minima are common for deep neural networks with piecewise linear activations
Liu, B. (2021) · 2021
Later among the works it cites.
Noise and fluctuation of finite learning rate stochastic gradient descent
Liu, K., Ziyin, L., and Ueda, M. (2021) · 2021
Later among the works it cites.
The loss surface of deep linear networks viewed through the algebraic geometry lens
Mehta, D., Chen, T., Tang, T., and Hauenstein, J. (2021) · 2021
Later among the works it cites.
Sgd can converge to local maxima
Ziyin, L., Li, B., Simon, J. B., and Ueda, M. (2021) · 2021
Later among the works it cites.
Power-law escape rate of sgd
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation
Gotmare, A., Keskar, N. S., Xiong, C., and Socher, R. (2018) · 2018
Cited alongside, same era.
An alternative view: When does sgd escape local minima?
Kleinberg, B., Li, Y., and Yuan, Y. (2018) · 2018
Cited alongside, same era.
Deep linear networks with arbitrary loss: All local minima are global
Laurent, T. and Brecht, J. (2018) · 2018
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
Safran, I. and Shamir, O. (2018) · 2018
Cited alongside, same era.
Small nonlinearities in activation functions create bad local minima in neural networks
Yun, C., Sra, S., and Jadbabaie, A. (2018) · 2018
Cited alongside, same era.
Don’t Blame the ELBO! A Linear VAE Perspective on Posterior Collapse
Lucas, J., Tucker, G., Grosse, R., and Norouzi, M. (2019) · 2019
Cited alongside, same era.
Mori, T., Ziyin, L., Liu, K., and Ueda, M. (2022) · 2022
Closest in time.
Neural collapse in deep homogeneous classifiers and the role of weight decay
Rangamani, A. and Banburski-Fahey, A. (2022) · 2022
Closest in time.
Deep contrastive learning is provably (almost) principal component analysis
Tian, Y. (2022) · 2022
Closest in time.
Posterior collapse of a linear latent variable model
Wang, Z. and Ziyin, L. (2022) · 2022
Closest in time.
Exact phase transitions in deep learning
Ziyin, L. and Ueda, M. (2022) · 2022
Closest in time.
Stochastic neural networks with infinite width are deterministic
Ziyin, L., Zhang, H., Meng, X., Lu, Y., Xing, E., and Ueda, M. (2022) · 2022
Closest in time.
What shapes the loss landscape of self supervised learning?
Ziyin, L., Lubana, E. S., Ueda, M., and Tanaka, H. (2023) · 2023
Closest in time.
spred: Solving L1 Penalty with SGD
Ziyin, L. and Wang, Z. (2023) · 2023
Closest in time.