Fetching the paper…
Reading the bibliography…
Motivated by the super-diffusivity of self-repelling random walk, which has roots in statistical physics, this paper develops a new perturbation mechanism for optimization algorithms.
When does non-orthogonal tensor decomposition have no spurious local minima?
Sanjabi, M., Baharlouei, S., Razaviyayn, M., and Lee, J. D · 1911
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
Optimization by simulated annealing
Kirkpatrick, S., Gelatt, J. C. D., and Vecchi, M. P · 1983
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Nesterov, Y. E · 1983
Earlier work this paper cites.
Cube slicing in 𝐑 n {\bf R}^{n}
Ball, K · 1986
Earlier work this paper cites.
Diffusions for global optimization
Geman, S. and Hwang, C.-R · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Random walks with memory
Peliti, L. and Pietronero, L · 1987
Earlier work this paper cites.
Asymptotics of the spectral gap with applications to the theory of simulated annealing
Holley, R. A., Kusuoka, S., and Stroock, D. W · 1989
Earlier work this paper cites.
Reinforced random walk
Davis, B · 1990
Earlier work this paper cites.
Recuit simulé sur 𝐑 n {\bf R}^{n} . Étude de l’évolution de l’énergie libre
Miclo, L · 1992
Earlier work this paper cites.
Vertex-reinforced random walk
Pemantle, R · 1992
Earlier work this paper cites.
The “true” self-avoiding walk with bond repulsion on ℤ \mathbb{Z} : limit theorems
Tóth, B · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Avoiding spurious local minima in deep quadratic networks
Kazemipour, A., Larsen, B., and Druckmann, S · 2001
Earlier work this paper cites.
Non-convex optimization via non-reversible stochastic gradient Langevin dynamics
Hu, Y., Wang, X., Gao, X., Gürbüzbalaban, M., and Zhu, L · 2004
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87 of Applied Optimization
Nesterov, Y · 2004
Earlier work this paper cites.
Vertex-reinforced random walk on ℤ \mathbb{Z} eventually gets stuck on five points
Tarrès, P · 2004
Earlier work this paper cites.
Phase transition in vertex-reinforced random walks on ℤ \mathbb{Z} with non-linear reinforcement
Volkov, S · 2006
Earlier work this paper cites.
Lévy flights, non-local search and simulated annealing
Pavlyukevich, I · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Cuckoo search via Lévy flights
Yang, X.-S. and Deb, S · 2009
Earlier work this paper cites.
Robust principal component analysis?
Candès, E. J., Li, X., Ma, Y., and Wright, J · 2011
Cited alongside, same era.
Phase retrieval via matrix completion
Candès, E. J., Eldar, Y. C., Strohmer, T., and Voroninski, V · 2013
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Bubeck, S · 2015
Cited alongside, same era.
Phase retrieval via Wirtinger flow: theory and algorithms
Candès, E. J., Li, X., and Soltanolkotabi, M · 2015
Cited alongside, same era.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Ben Arous, G., and LeCun, Y · 2015
Cited alongside, same era.
Escaping from saddle points – online stochastic gradient for tensor decomposition
Ge, R., Huang, F., Jin, C., and Yuan, Y · 2015
Cited alongside, same era.
Non-convex learning via stochastic gradient Langevin dynamics: a nonasymptotic analysis
Raginsky, M., Rakhlin, A., and Telgarsky, M · 2017
Later among the works it cites.
Complete dictionary recovery over the sphere I: Overview and the geometric picture
Sun, J., Qu, Q., and Wright, J · 2017
Later among the works it cites.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F · 2018
Later among the works it cites.
Gradient descent learns one-hidden-layer CNN: Don’t be afraid of spurious local minima
Du, S. S., Lee, J. D., Tian, Y., Singh, A., and Poczos, B · 2018
Later among the works it cites.
Learning one-hidden-layer neural networks with landscape design
Ge, R., Lee, J. D., and Ma, T · 2018
Later among the works it cites.
Accelerated gradient descent escapes saddle points faster than gradient descent
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Adding gradient noise improves learning for very deep networks
Neelakantan, A., Vilnis, L., Le, Q. V., Sutskever, I., Kaiser, L., Kurach, K., and Martens, J · 2015
Cited alongside, same era.
On the low-rank approach for semidefinite programs arising in synchronization and community detection
Bandeira, A. S., Boumal, N., and Voroninski, V · 2016
Cited alongside, same era.
Global optimality of local search for low rank matrix recovery
Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2016
Cited alongside, same era.
Matrix completion has no spurious local minimum
Ge, R., Lee, J. D., and Ma, T · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K · 2016
Cited alongside, same era.
Jin, C., Netrapalli, P., and Jordan, M. I · 2018
Later among the works it cites.
Adding one neuron can eliminate all bad local minima
Liang, S., Sun, R., Lee, J. D., and Srikant, R · 2018
Later among the works it cites.
The landscape of empirical risk for nonconvex losses
Mei, S., Bai, Y., and Montanari, A · 2018
Later among the works it cites.
Ergodicity of the infinite swapping algorithm at low temperature
Menz, G., Schlichting, A., Tang, W., and Wu, T · 2018
Later among the works it cites.
Hypocoercivity in metastable settings and kinetic simulated annealing
Monmarché, P · 2018
Later among the works it cites.
A geometric analysis of phase retrieval
Sun, J., Qu, Q., and Wright, J · 2018
Later among the works it cites.
No spurious local minima in a two hidden unit ReLU network
Wu, C., Luo, J., and Lee, J. D · 2018
Later among the works it cites.
Accelerating nonconvex learning via replica exchange Langevin diffusion
Chen, Y., Chen, J., Dong, J., Peng, J., and Wang, Z · 2019
Later among the works it cites.
Spurious valleys in one-hidden-layer neural network optimization landscapes
Venturi, L., Bandeira, A. S., and Bruna, J · 2019
Later among the works it cites.
Escaping saddle points faster with stochastic momentum
Wang, J.-K., Lin, C.-H., and Abernethy, J · 2019
Later among the works it cites.
Toward understanding the importance of noise in training neural networks
Zhou, M., Liu, T., Li, Y., Lin, D., Zhou, E., and Zhao, T · 2019
Later among the works it cites.
On stationary-point hitting time and ergodicity of stochastic gradient Langevin dynamics
Chen, X., Du, S. S., and Tong, X. T · 2020
Closest in time.
Replica exchange for non-convex optimization
Dong, J. and Tong, X. T · 2021
Closest in time.
On nonconvex optimization for machine learning: Gradients, stochasticity, and saddle points
Jin, C., Netrapalli, P., Ge, R., Kakade, S. M., and Jordan, M. I · 2021
Closest in time.
On the curse of memory in recurrent neural networks: Approximation and optimization analysis
Li, Z., Han, J., E, W., and Li, Q · 2021
Closest in time.
Tang, W. and Zhou, X. Y · 2021
Closest in time.