Fetching the paper…
Reading the bibliography…
In overparametrized models, the noise in stochastic gradient descent (SGD) implicitly regularizes the optimization trajectory and determines which local minimum SGD converges to.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Efficient backprop
Y. A. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 2012
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2017
Earlier work this paper cites.
Y. Li, T. Ma, and H. Zhang · 2017
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
S. L. Smith, P.-J. Kindermans, C. Ying, and Q. V. Le · 2017
Earlier work this paper cites.
Characterizing implicit bias in terms of optimization geometry
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Earlier work this paper cites.
On the implicit bias of dropout
P. Mianjy, R. Arora, and R. Vidal · 2018
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro · 2018
Cited alongside, same era.
On exact computation with an infinitely wide neural net
S. Arora, S. S. Du, W. Hu, Z. Li, R. Salakhutdinov, and R. Wang · 2019
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Implicit regularization for optimal sparse recovery
T. Vaskevicius, V. Kanade, and P. Rebeschini · 2019
Later among the works it cites.
Improved sample complexities for deep networks and robust classification via an all-layer margin
C. Wei and T. Ma · 2019
Later among the works it cites.
Y. Wen, K. Luk, M. Gazeau, G. Zhang, H. Chan, and J. Ba · 2019
Later among the works it cites.
Experiment tracking with weights and biases, 2020
L. Biewald · 2020
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Blanc, N. Gupta, G. Valiant, and P. Valiant · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks, 2019
S. S. Du, J. D. Lee, H. Li, L. Wang, and X. Zhai · 2019
Cited alongside, same era.
Pytorch lightning
W. Falcon et al · 2019
Cited alongside, same era.
Stochastic gradient descent escapes saddle points efficiently
C. Jin, P. Netrapalli, R. Ge, S. M. Kakade, and M. I. Jordan · 2019
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Y. Li, C. Wei, and T. Ma · 2019
Cited alongside, same era.
Bad global minima exist and sgd can reach them
S. Liu, D. Papailiopoulos, and D. Achlioptas · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Cited alongside, same era.
Shape matters: Understanding the implicit bias of the noise covariance
J. Z. HaoChen, C. Wei, J. D. Lee, and T. Ma · 2020
Later among the works it cites.
Gaussian error linear units (gelus), 2020
D. Hendrycks and K. Gimpel · 2020
Later among the works it cites.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
E. Moroshko, S. Gunasekar, B. Woodworth, J. D. Lee, N. Srebro, and D. Soudry · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
B. Woodworth, S. Gunasekar, J. D. Lee, E. Moroshko, P. Savarese, I. Golan, D. Soudry, and N. Srebro · 2020
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability, 2021
J. M. Cohen, S. Kaur, Y. Li, J. Z. Kolter, and A. Talwalkar · 2021
Closest in time.