Hessian-based analysis of large batch training and robustness to adversaries
Z. Yao, A. Gholami, Q. Lei, K. Keutzer, and M. W. Mahoney · 2018
Later among the works it cites.
Removing the feature correlation effect of multiplicative noise
Z. Zhang, Y. Zhang, and Z. Li · 2018
Later among the works it cites.
On the convergence of adaptive gradient methods for nonconvex optimization
Original
D. Zhou, Y. Tang, Z. Yang, Y. Cao, and Q. Gu · 2018
Later among the works it cites.
On the convergence of a class of Adam-type algorithms for non-convex optimization
X. Chen, S. Liu, R. Sun, and M. Hong · 2019
Later among the works it cites.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2019
Later among the works it cites.
SGD: General analysis and improved rates
R. M. Gower, N. Loizou, X. Qian, A. Sailanbayev, E. Shulgin, and P. Richtárik · 2019
Later among the works it cites.
Products of many large random matrices and gradients in deep neural networks
B. Hanin and M. Nica · 2019
Later among the works it cites.
Robust descent using smoothed multiplicative noise
M. J. Holland · 2019
Later among the works it cites.
Deep learning theory review: An optimal control and dynamical systems perspective
Original
G.-H. Liu and E. A. Theodorou · 2019
Later among the works it cites.
Bad global minima exist and SGD can reach them
Original
S. Liu, D. Papailiopoulos, and D. Achlioptas · 2019
Later among the works it cites.
Traditional and heavy-tailed self regularization in neural network models
C. H. Martin and M. W. Mahoney · 2019
Later among the works it cites.
SGD without replacement: Sharper rates for general smooth convex functions
D. Nagaraj, P. Netrapalli, and P. Jain · 2019
Later among the works it cites.
Continuous-time models for stochastic optimization algorithms
A. Orvieto and A. Lucchi · 2019
Later among the works it cites.
Generalization, Adaptation and Low-Rank Representation in Neural Networks
S. Oymak, Z. Fabian, M. Li, and M. Soltanolkotabi · 2019
Later among the works it cites.
Non-gaussianity of stochastic gradient noise
A. Panigrahi, R. Somani, N. Goyal, and P. Netrapalli · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
U. Şimşekli, L. Sagun, and M. Gürbüzbalaban · 2019
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
R. Ward, X. Wu, and L. Bottou · 2019
Later among the works it cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
G. Zhang, L. Li, Z. Nado, J. Martens, S. Sachdeva, G. Dahl, C. Shallue, and R. B. Grosse · 2019
Later among the works it cites.
Continuous and Discrete-Time Analysis of Stochastic Gradient Descent for Convex and Non-Convex Functions
Original
X. Fontaine, V. De Bortoli, and A. Durmus · 2020
Closest in time.
The Heavy-Tail Phenomenon in SGD
Original
M. Gürbüzbalaban, U. Şimşekli, and L. Zhu · 2020
Closest in time.
The reproducing Stein kernel approach for post-hoc corrected sampling
Original
L. Hodgkinson, R. Salomone, and F. Roosta · 2020
Closest in time.
Budgeted Training: Rethinking Deep Neural Network Training Under Resource Constraints
M. Li, E. Yumer, and D. Ramanan · 2020
Closest in time.
On the variance of the adaptive learning rate and beyond
L. Liu, H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and J. Han · 2020
Closest in time.
Heavy-tailed Universality predicts trends in test accuracies for very large pre-trained deep neural networks
C. H. Martin and M. W. Mahoney · 2020
Closest in time.
Predicting trends in the quality of state-of-the-art neural networks without access to training or testing data
Original
C. H. Martin and M. W. Mahoney · 2020
Closest in time.
Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient Noise
Original
U. Şimşekli, L. Zhu, Y. W. Teh, and M. Gürbüzbalaban · 2020
Closest in time.
How good is the bayes posterior in deep neural networks really?
Original
F. Wenzel, K. Roth, B. S. Veeling, J. Świątkowski, L. Tran, S. Mandt, J. Snoek, T. Salimans, R. Jenatton, and S. Nowozin · 2020
Closest in time.
On the Noisy Gradient Descent that Generalizes as SGD
Original
J. Wu, W. Hu, H. Xiong, J. Huan, V. Braverman, and Z. Zhu · 2020
Closest in time.