Fetching the paper…
Reading the bibliography…
In this work, we study the evolution of the loss Hessian across many classification tasks in order to understand the effect the curvature of the loss has on the training dynamics.
Fixup initialization: Residual learning without normalization
H. Zhang, Y. N. Dauphin, and T. Ma · 1901
Earlier work this paper cites.
Fast exact multiplication by the hessian
B. A. Pearlmutter · 1994
Earlier work this paper cites.
Ecological paradoxes: William stanley jevons and the paperless office
R. York · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, and P. Koehn · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Earlier work this paper cites.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
S. Jastrzębski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2017
Cited alongside, same era.
Visualizing the loss landscape of neural nets
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2017
Cited alongside, same era.
An empirical model of large-batch training
S. McCandlish, J. Kaplan, D. Amodei, and O. D. Team · 2018
A mean field theory of batch normalization
G. Yang, J. Pennington, V. Rao, J. Sohl-Dickstein, and S. S. Schoenholz · 2019
Later among the works it cites.
The break-even point on optimization trajectories of deep neural networks
S. Jastrzebski, M. Szymczak, S. Fort, D. Arpit, J. Tabor, K. Cho, and K. Geras · 2020
Later among the works it cites.
The large learning rate phase of deep learning: the catapult mechanism
A. Lewkowycz, Y. Bahri, E. Dyer, J. Sohl-Dickstein, and G. Gur-Ari · 2020
Later among the works it cites.
Understanding the difficulty of training transformers
L. Liu, X. Liu, J. Gao, W. Chen, and J. Han · 2020
Later among the works it cites.
Optimizer benchmarking needs to account for hyperparameter tuning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size
V. Papyan · 2018
Cited alongside, same era.
How does batch normalization help optimization?
S. Santurkar, D. Tsipras, A. Ilyas, and A. Madry · 2018
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl · 2018
Cited alongside, same era.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
L. Wu, C. Ma, and W. E · 2018
Cited alongside, same era.
On empirical comparisons of optimizers for deep learning
D. Choi, C. J. Shallue, Z. Nado, J. Lee, C. J. Maddison, and G. E. Dahl · 2019
Cited alongside, same era.
Metainit: Initializing learning by learning to initialize
Y. N. Dauphin and S. Schoenholz · 2019
Cited alongside, same era.
An investigation into neural net optimization via hessian eigenvalue density
B. Ghorbani, S. Krishnan, and Y. Xiao · 2019
Cited alongside, same era.
P. T. Sivaprasad, F. Mai, T. Vogels, M. Jaggi, and F. Fleuret · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T. Liu · 2020
Later among the works it cites.
Characterizing signal propagation to close the performance gap in unnormalized resnets
A. Brock, S. De, and S. L. Smith · 2021
Closest in time.
Gradient descent on neural networks typically occurs at the edge of stability
J. M. Cohen, S. Kaur, Y. Li, J. Z. Kolter, and A. Talwalkar · 2021
Closest in time.
A large batch optimizer reality check: Traditional, generic optimizers suffice across batch sizes
Z. Nado, J. Gilmer, C. J. Shallue, R. Anil, and G. E. Dahl · 2021
Closest in time.
Carbon emissions and large neural network training
D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean · 2021
Closest in time.
Gradinit: Learning to initialize neural networks for stable and efficient training
C. Zhu, R. Ni, Z. Xu, K. Kong, W. R. Huang, and T. Goldstein · 2021
Closest in time.