Fetching the paper…
Reading the bibliography…
Mode connectivity is a recently introduced frame- work that empirically establishes the connected- ness of minima by finding a high accuracy curve between two independently trained models.
Flat minima
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Visualizing data using t-sne
Maaten, Laurens van der and Hinton, Geoffrey · 2008
Earlier work this paper cites.
The marginal value of adaptive gradient methods in machine learning
Wilson, Ashia C, Roelofs, Rebecca, Stern, Mitchell, Srebro, Nati, and Recht, Benjamin · 2008
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, Christian, Zaremba, Wojciech, Sutskever, Ilya, Bruna, Joan, Erhan, Dumitru, Goodfellow, Ian, and Fergus, Rob · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, Diederik P and Ba, Jimmy · 2014
Cited alongside, same era.
The cifar-10 dataset
Krizhevsky, Alex, Nair, Vinod, and Hinton, Geoffrey · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, Karen and Zisserman, Andrew · 2014
Cited alongside, same era.
Deep learning , volume 1
Goodfellow, Ian, Bengio, Yoshua, Courville, Aaron, and Bengio, Yoshua · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, Nitish Shirish, Mudigere, Dheevatsa, Nocedal, Jorge, Smelyanskiy, Mikhail, and Tang, Ping Tak Peter · 2016
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a nash equilibrium
Heusel, Martin, Ramsauer, Hubert, Unterthiner, Thomas, Nessler, Bernhard, Klambauer, Günter, and Hochreiter, Sepp · 2017
Later among the works it cites.
Snapshot ensembles: Train 1, get m for free
Huang, Gao, Li, Yixuan, Pleiss, Geoff, Liu, Zhuang, Hopcroft, John E, and Weinberger, Kilian Q · 2017
Later among the works it cites.
Cyclical learning rates for training neural networks
Smith, Leslie N · 2017
Later among the works it cites.
Essentially no barriers in neural network energy landscape
Draxler, Felix, Veschgini, Kambis, Salmhofer, Manfred, and Hamprecht, Fred A · 2018
Closest in time.
Loss surfaces, mode connectivity, and fast ensembling of dnns
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Loshchilov, Ilya and Hutter, Frank · 2016
Cited alongside, same era.
Dawnbench: An end-to-end deep learning benchmark and competition
Coleman, Cody, Narayanan, Deepak, Kang, Daniel, Zhao, Tian, Zhang, Jian, Nardi, Luigi, Bailis, Peter, Olukotun, Kunle, Ré, Chris, and Zaharia, Matei · 2017
Cited alongside, same era.
Garipov, Timur, Izmailov, Pavel, Podoprikhin, Dmitrii, Vetrov, Dmitry P, and Wilson, Andrew Gordon · 2018
Closest in time.
Averaging weights leads to wider optima and better generalization
Izmailov, Pavel, Podoprikhin, Dmitrii, Garipov, Timur, Vetrov, Dmitry, and Wilson, Andrew Gordon · 2018
Closest in time.