Fetching the paper…
Reading the bibliography…
There are many surprising and perhaps counter-intuitive properties of optimization of deep neural networks.
Efficient backprop
Yann LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems, 2014
Ian J. Goodfellow, Oriol Vinyals, and Andrew M. Saxe · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gerard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Singularity of the hessian in deep learning
Levent Sagun, Léon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
Three factors influencing minima in sgd, 2017
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
The loss surface of deep and wide neural networks
Quynh Nguyen and Matthias Hein · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Cited alongside, same era.
Snapshot ensembles: Train 1, get m for free, 2017
Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E. Hopcroft, and Kilian Q. Weinberger · 2017
Cited alongside, same era.
The goldilocks zone: Towards better understanding of neural network loss landscapes, 2018
Stanislav Fort and Adam Scherlis · 2018
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L. Smith, Pieter-Jan Kindermans, and Quoc V. Le · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization, 2018
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Measuring the intrinsic dimension of objective landscapes, 2018a
On the Relation Between the Sharpest Directions of DNN Loss and the SGD Step Length
Stanisław Jastrzębski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Later among the works it cites.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Later among the works it cites.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred A Hamprecht · 2018
Later among the works it cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski
Cited in the paper.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein
Cited in the paper.
Closest in time.