Fetching the paper…
Reading the bibliography…
Neural network training is commonly based on SGD.
Differentiable dynamical systems
Smale, S · 1967
Earlier work this paper cites.
Framed Morse complexes and its invariants
Barannikov, S · 1994
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Geometry helps in bottleneck matching and related problems
Efrat, A., Itai, A., and Katz, M. J · 2001
Earlier work this paper cites.
Computing and comprehending topology: Persistence and hierarchical Morse complexes (Ph.D.Thesis)
Zomorodian, A. J · 2001
Earlier work this paper cites.
Precise Arrhenius law for p-forms: The Witten Laplacian and Morse–Barannikov complex
Le Peutrec, D., Nier, F., and Viterbo, C · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Earlier work this paper cites.
Lectures on the h-cobordism theorem
Milnor, J · 2015
Earlier work this paper cites.
Cyclical learning rates for training neural networks, 2015
Smith, L. N · 2015
Cited alongside, same era.
Deep learning , volume 1
Goodfellow, I., Bengio, Y., Courville, A., and Bengio, Y · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Cited alongside, same era.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
An introduction to topological data analysis: fundamental and practical aspects for data scientists
Chazal, F. and Michel, B · 2017
Cited alongside, same era.
A deep neural network’s loss surface contains every low-dimensional pattern, 2019
Czarnecki, W. M., Osindero, S., Pascanu, R., and Jaderberg, M · 2019
Later among the works it cites.
Large scale structure of neural network loss landscapes
Fort, S. and Jastrzebski, S · 2019
Later among the works it cites.
Asymmetric valleys: Beyond sharp and flat local minima
He, H., Huang, G., and Yuan, Y · 2019
Later among the works it cites.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Kuditipudi, R., Wang, X., Lee, H., Zhang, Y., Li, Z., Hu, W., Ge, R., and Arora, S · 2019
Later among the works it cites.
Li, Y., Wei, C., and Ma, T · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Cited alongside, same era.
A hitting time analysis of stochastic gradient langevin dynamics
Zhang, Y., Liang, P., and Charikar, M · 2017
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. A · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Cited alongside, same era.
Using mode connectivity for loss landscape analysis
Gotmare, A., Keskar, N. S., Xiong, C., and Socher, R · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Cited alongside, same era.
Barcodes as summary of objective function’s topology
Barannikov, S., Korotin, A., Oganesyan, D., Emtsev, D., and Burnaev, E · 2019
Cited alongside, same era.
Nguyen, Q · 2019
Later among the works it cites.
Loss landscape sightseeing with multi-point optimization
Skorokhodov, I. and Burtsev, M · 2019
Later among the works it cites.
Pllay: Efficient topological layer based on persistence landscapes
Kim, K., Kim, J., Zaheer, M., Kim, J., Chazal, F., and Wasserman, L · 2020
Closest in time.
Topological autoencoders
Moor, M., Horn, M., Rieck, B., and Borgwardt, K · 2020
Closest in time.
On learning rates and Schrödinger operators
Shi, B., Su, W. J., and Jordan, M. I · 2020
Closest in time.
Bridging mode connectivity in loss landscapes and adversarial robustness
Zhao, P., Chen, P.-Y., Das, P., Ramamurthy, K. N., and Lin, X · 2020
Closest in time.
Canonical Forms = Persistence Diagrams. Tutorial
Barannikov, S · 2021
Closest in time.