“Large Scale Structure of Neural Network Loss Landscapes”, 2019
Original
Stanislav Fort and Stanislaw Jastrzebski · 1906
Earlier work this paper cites.
“Emergent properties of the local geometry of neural loss landscapes”, 2019
Original
Stanislav Fort and Surya Ganguli · 1910
Earlier work this paper cites.
“Linear Mode Connectivity and the Lottery Ticket Hypothesis”
Original
Jonathan Frankle, Gintare Dziugaite, Daniel Roy and Michael Carbin · 1912
Earlier work this paper cites.
“Deep Ensembles: A Loss Landscape Perspective”, 2019
Original
Stanislav Fort, Huiyi Hu and Balaji Lakshminarayanan · 1912
Earlier work this paper cites.
“Neural Tangents: Fast and Easy Infinite Neural Networks in Python”, 2019
Original
Roman Novak et al · 1912
Earlier work this paper cites.
“Flat minima”
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
“The Break-Even Point on Optimization Trajectories of Deep Neural Networks”, 2020
Original
Stanislaw Jastrzebski et al · 2002
Earlier work this paper cites.
“The Two Regimes of Deep Network Training”, 2020
Original
Guillaume Leclerc and Aleksander Madry · 2002
Earlier work this paper cites.
“Taylorized Training: Towards Better Approximation of Neural Network Training at Finite Width”, 2020
Original
Yu Bai et al · 2002
Earlier work this paper cites.
“Bounding the True Error”
John Langford and Rich Caruana · 2002
Earlier work this paper cites.
“The large learning rate phase of deep learning: the catapult mechanism”, 2020
Original
Aitor Lewkowycz et al · 2003
Earlier work this paper cites.
“Visualizing Data using t-SNE”
Laurens van Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
“Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification”, 2015
Original
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2015
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”, 2015
Original
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2015
Earlier work this paper cites.