Fetching the paper…
Reading the bibliography…
Numerous empirical evidences have corroborated the importance of noise in nonconvex optimization problems.
On connected sublevel sets in deep learning
Nguyen, Q · 1901
Earlier work this paper cites.
Towards understanding the importance of noise in training neural networks
Zhou, M · 1909
Earlier work this paper cites.
A variational baysian framework for graphical models
Attias, H · 2000
Earlier work this paper cites.
Graphical models
Jordan, M. I · 2004
Earlier work this paper cites.
Shape matters: Understanding the implicit bias of the noise covariance
HaoChen, J. Z · 2006
Earlier work this paper cites.
Restricted boltzmann machines for collaborative filtering
Salakhutdinov, R · 2007
Earlier work this paper cites.
Matrix completion from a few entries
Keshavan, R. H · 2010
Earlier work this paper cites.
Randomized smoothing for stochastic optimization
Duchi, J. C · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A · 2012
Earlier work this paper cites.
Understanding alternating minimization for matrix completion
Hardt, M · 2014
Earlier work this paper cites.
Chen, Y · 2015
Earlier work this paper cites.
Low-rank solutions of linear matrix equations via procrustes flow
Tu, S · 2015
Cited alongside, same era.
A nonconvex optimization framework for low rank matrix estimation
Zhao, T · 2015
Cited alongside, same era.
Dropping convexity for faster semi-definite optimization
Bhojanapalli, S · 2016
Cited alongside, same era.
Matrix completion has no spurious local minimum
Ge, R · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S · 2016
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Du, S. S · 2018
Later among the works it cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T · 2018
Later among the works it cites.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S · 2018
Later among the works it cites.
Gradient descent aligns the layers of deep linear networks
Ji, Z · 2018
Later among the works it cites.
Understanding the loss surface of neural networks for binary classification
Liang, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Zheng, Q · 2016
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Ge, R · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Jin, C · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Nguyen, Q · 2017
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Draxler, F · 2018
Cited alongside, same era.
Luo, P · 2018
Later among the works it cites.
The loss surface and expressivity of deep convolutional neural networks
Nguyen, Q · 2018
Later among the works it cites.
Spurious valleys in two-layer neural network optimization landscapes
Venturi, L · 2018
Later among the works it cites.
Nonconvex optimization meets low-rank matrix factorization: An overview
Chi, Y · 2019
Later among the works it cites.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Kuditipudi, R · 2019
Later among the works it cites.
First-order methods almost always avoid strict saddle points
Lee, J. D · 2019
Later among the works it cites.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Blanc, G · 2020
Later among the works it cites.