Fetching the paper…
Reading the bibliography…
It is widely believed that the implicit regularization of SGD is fundamental to the impressive generalization behavior we observe in neural networks.
Large-Batch Training for LSTM and Beyond
Yang You, Jonathan Hseu, Chris Ying, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 1901
Earlier work this paper cites.
On Nonconvex Optimization for Machine Learning: Gradients, Stochasticity, and Saddle Points
Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M. Kakade, and Michael I. Jordan · 1902
Earlier work this paper cites.
Yet Another Accelerated SGD: ResNet-50 Training on ImageNet in 74.7 seconds
Masafumi Yamazaki, Akihiko Kasagi, Akihiro Tabuchi, Takumi Honda, Masahiro Miwa, Naoto Fukumoto, Tsuguchika Tabaru, Atsushi Ike, and Kohta Nakashima · 1903
Earlier work this paper cites.
Implicit Regularization in Deep Matrix Factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 1905
Earlier work this paper cites.
Understanding Generalization through Visualizations
W. Ronny Huang, Zeyad Emam, Micah Goldblum, Liam Fowl, Justin K. Terry, Furong Huang, and Tom Goldstein · 1906
Earlier work this paper cites.
Yuanzhi Li, Colin Wei, and Tengyu Ma · 1907
Earlier work this paper cites.
Fantastic Generalization Measures and Where to Find Them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 1912
Earlier work this paper cites.
Fast Exact Multiplication by the Hessian
Barak A. Pearlmutter · 1994
Earlier work this paper cites.
Flat Minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Scalable Second Order Optimization for Deep Learning
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer · 2002
Earlier work this paper cites.
Can Implicit Bias Explain Generalization? Stochastic Convex Optimization as a Case Study
Assaf Dauber, Meir Feder, Tomer Koren, and Roi Livni · 2003
Earlier work this paper cites.
The large learning rate phase of deep learning: The catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2003
Earlier work this paper cites.
The general inefficiency of batch training for gradient descent learning
D. Randall Wilson and Tony R. Martinez · 2003
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Shape Matters: Understanding the Implicit Bias of the Noise Covariance
Jeff Z. HaoChen, Colin Wei, Jason D. Lee, and Tengyu Ma · 2006
Earlier work this paper cites.
Yang You, Yuhui Wang, Huan Zhang, Zhao Zhang, James Demmel, and Cho-Jui Hsieh · 2006
Earlier work this paper cites.
Descending through a Crowded Valley – Benchmarking Deep Learning Optimizers
Robin M. Schmidt, Frank Schneider, and Philipp Hennig · 2007
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Large-Scale Machine Learning with Stochastic Gradient Descent
Léon Bottou · 2010
Earlier work this paper cites.
Sharpness-Aware Minimization for Efficiently Improving Generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2010
Earlier work this paper cites.
Are wider nets better given the same number of parameters?
Anna Golubeva, Behnam Neyshabur, and Guy Gur-Ari · 2010
Earlier work this paper cites.
Understanding and Scheduling Weight Decay
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
On the difficulty of training Recurrent Neural Networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
cuDNN: Efficient Primitives for Deep Learning
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer · 2014
Earlier work this paper cites.
Escaping From Saddle Points — Online Stochastic Gradient for Tensor Decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Gradient Descent Only Converges to Minimizers
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Earlier work this paper cites.
Automated Inference with Adaptive Batches
Soham De, Abhay Yadav, David Jacobs, and Tom Goldstein · 2017
Earlier work this paper cites.
Sharp Minima Can Generalize For Deep Nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Gradient Descent Can Take Exponential Time to Escape Saddle Points
Simon S. Du, Chi Jin, Jason D. Lee, Michael I. Jordan, Aarti Singh, and Barnabas Poczos · 2017
Earlier work this paper cites.
Machine learning - Why mini batch size is better than one single "batch" with all training data?, February 2017
Hendrik · 2017
Cited alongside, same era.
Train longer, generalize better: Closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Stochastic Modified Equations and Adaptive Stochastic Gradient Algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Cited alongside, same era.
SGDR: Stochastic Gradient Descent with Warm Restarts
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Bag of Tricks for Image Classification with Convolutional Neural Networks
Tong He, Zhi Zhang, Hang Zhang, Zhongyue Zhang, Junyuan Xie, and Mu Li · 2019
Later among the works it cites.
An Exponential Learning Rate Schedule for Deep Learning
Zhiyuan Li and Sanjeev Arora · 2019
Later among the works it cites.
Massively Distributed SGD: ImageNet/ResNet-50 Training in a Flash
Hiroaki Mikami, Hisahiro Suganuma, Pongsakorn U-chupala, Yoshiki Tanaka, and Yuichi Kageyama · 2019
Later among the works it cites.
Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2019
Later among the works it cites.
On the information bottleneck theory of deep learning
Andrew M. Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D. Tracey, and David D. Cox · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lei Wu, Zhanxing Zhu, and Weinan E · 2017
Cited alongside, same era.
Large Batch Training of Convolutional Networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Cited alongside, same era.
Towards Theoretical Understanding of Large Batch Training in Stochastic Gradient Descent
Xiaowu Dai and Yuhua Zhu · 2018
Cited alongside, same era.
On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent
Noah Golmant, Nikita Vemuri, Zhewei Yao, Vladimir Feinberg, Amir Gholami, Kai Rothauge, Michael Mahoney, and Joseph Gonzalez · 2018
Cited alongside, same era.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2018
Cited alongside, same era.
Measuring the Effects of Data Parallelism on Neural Network Training
Christopher J. Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl · 2019
Later among the works it cites.
A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
Later among the works it cites.
Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George Dahl, Chris Shallue, and Roger B. Grosse · 2019
Later among the works it cites.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Later among the works it cites.
High-dimensional dynamics of generalization error in neural networks
Madhu S. Advani, Andrew M. Saxe, and Haim Sompolinsky · 2020
Later among the works it cites.
Implicit Gradient Regularization
David Barrett and Benoit Dherin · 2020
Later among the works it cites.
Stability of stochastic gradient descent on nonsmooth convex losses
Raef Bassily, Vitaly Feldman, Cristóbal Guzmán, and Kunal Talwar · 2020
Later among the works it cites.
Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based Optimization
Satrajit Chatterjee · 2020
Later among the works it cites.
Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet Talwalkar · 2020
Later among the works it cites.
Truth or backpropaganda? An empirical investigation of deep learning theory
Micah Goldblum, Jonas Geiping, Avi Schwarzschild, Michael Moeller, and Tom Goldstein · 2020
Later among the works it cites.
Stochastic optimization with heavy-tailed noise via accelerated gradient clipping
Eduard Gorbunov, Marina Danilova, and Alexander Gasnikov · 2020
Later among the works it cites.
Stochasticity of Deterministic Gradient Descent: Large Learning Rate for Multiscale Objective Function
Lingkai Kong and Molei Tao · 2020
Later among the works it cites.
Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning Rate
Zhiyuan Li, Kaifeng Lyu, and Sanjeev Arora · 2020
Later among the works it cites.
Optimizing Neural Networks with Kronecker-factored Approximate Curvature
James Martens and Roger Grosse · 2020
Later among the works it cites.
Extreme Memorization via Scale of Initialization
Harsh Mehta, Ashok Cutkosky, and Behnam Neyshabur · 2020
Later among the works it cites.
Loss landscape: Sgd has a better view
Tomaso Poggio and Yaim Cooper · 2020
Later among the works it cites.
An Empirical Study of Stochastic Gradient Descent with Structured Covariance Noise
Yeming Wen, Kevin Luk, Maxime Gazeau, Guodong Zhang, Harris Chan, and Jimmy Ba · 2020
Later among the works it cites.
On the Noisy Gradient Descent that Generalizes as SGD
Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, and Zhanxing Zhu · 2020
Later among the works it cites.
A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2020
Later among the works it cites.
Making L-BFGS Work with Industrial-Strength Nets
Abhay Yadav · 2020
Later among the works it cites.
Improved Analysis of Clipping Algorithms for Non-convex Optimization
Bohang Zhang, Jikai Jin, Cong Fang, and Liwei Wang · 2020
Later among the works it cites.
Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep Learning
Pan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong, Steven Chu Hong Hoi, and Weinan E · 2020
Later among the works it cites.
Implicit Gradient Alignment in Distributed and Federated Learning
Yatin Dandi, Luis Barba, and Martin Jaggi · 2021
Closest in time.
A Loss Curvature Perspective on Training Instability in Deep Learning
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta, Behnam Neyshabur, David Cardoze, George Dahl, Zachary Nado, and Orhan Firat · 2021
Closest in time.
Daniel Kunin, Javier Sagastuy-Brena, Lauren Gillespie, Eshed Margalit, Hidenori Tanaka, Surya Ganguli, and Daniel L. K. Yamins · 2021
Closest in time.
On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)
Zhiyuan Li, Sadhika Malladi, and Sanjeev Arora · 2021
Closest in time.
NVIDIA cuDNN, 2022
NVIDIA · 2022
Closest in time.