Fetching the paper…
Reading the bibliography…
This article examines the implicit regularization effect of Stochastic Gradient Descent (SGD).
On Connected Sublevel Sets in Deep Learning
Quynh Nguyen · 1901
Earlier work this paper cites.
Sharp Analysis for Nonconvex SGD Escaping from Saddle Points, June 2019
Cong Fang, Zhouchen Lin, and Tong Zhang · 1902
Earlier work this paper cites.
Benign Overfitting in Linear Regression
Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler · 1906
Earlier work this paper cites.
Explaining Landscape Connectivity of Low-cost Solutions for Multilayer Nets
Rohith Kuditipudi, Xiang Wang, Holden Lee, Yi Zhang, Zhiyuan Li, Wei Hu, Sanjeev Arora, and Rong Ge · 1906
Earlier work this paper cites.
Fantastic Generalization Measures and Where to Find Them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 1912
Earlier work this paper cites.
Big Transfer (BiT): General Visual Representation Learning, May 2020
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby · 1912
Earlier work this paper cites.
Error analysis of floating-point computation
J. H. Wilkinson · 1960
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
The Break-Even Point on Optimization Trajectories of Deep Neural Networks
Stanisław Jastrzębski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2002
Earlier work this paper cites.
A Unified Convergence Analysis for Shuffling-Type Gradient Methods, September 2021
Lam M. Nguyen, Quoc Tran-Dinh, Dzung T. Phan, Phuong Ha Nguyen, and Marten van Dijk · 2002
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2003
Earlier work this paper cites.
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2003
Earlier work this paper cites.
Geometric numerical integration: structure-preserving algorithms for ordinary differential equations
E. Hairer, Christian Lubich, and Gerhard Wanner · 2006
Earlier work this paper cites.
Shape Matters: Understanding the Implicit Bias of the Noise Covariance
Jeff Z. HaoChen, Colin Wei, Jason D. Lee, and Tengyu Ma · 2006
Earlier work this paper cites.
Itay Safran, Gilad Yehudai, and Ohad Shamir · 2006
Earlier work this paper cites.
Traces of Class/Cross-Class Structure Pervade Deep Learning Spectra, August 2020
Vardan Papyan · 2008
Earlier work this paper cites.
Implicit Gradient Regularization
David G. T. Barrett and Benoit Dherin · 2009
Earlier work this paper cites.
Curiously Fast Convergence of some Stochastic Gradient Descent Algorithms
Leon Bottou · 2009
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures, September 2012
Yoshua Bengio · 2012
Earlier work this paper cites.
Stochastic Gradient Descent Tricks
Léon Bottou · 2012
Earlier work this paper cites.
Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts Generalization
Stanisław Jastrzębski, Devansh Arpit, Oliver Astrand, Giancarlo Kerg, Huan Wang, Caiming Xiong, Richard Socher, Kyunghyun Cho, and Krzysztof Geras · 2012
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Parallel stochastic gradient algorithms for large-scale matrix completion
Benjamin Recht and Christopher Ré · 2013
Earlier work this paper cites.
Yann Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems, May 2015
Ian J. Goodfellow, Oriol Vinyals, and Andrew M. Saxe · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Anima Anandkumar and Rong Ge · 2016
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Deep Learning without Poor Local Minima, December 2016
Kenji Kawaguchi · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Gradient Descent Only Converges to Minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
ON THE LOSS LANDSCAPE OF A CLASS OF DEEP NEURAL NETWORKS WITH NO BAD LOCAL VALLEYS
Quynh Nguyen, Matthias Hein, and Mahesh Chandra Mukkamala · 2019
Later among the works it cites.
Vardan Papyan · 2019
Later among the works it cites.
Yeming Wen, Kevin Luk, Maxime Gazeau, Guodong Zhang, Harris Chan, and Jimmy Ba · 2019
Later among the works it cites.
New Insights and Perspectives on the Natural Gradient Method
James Martens · 2020
Later among the works it cites.
Random Reshuffling: Simple Analysis with Vast Improvements
Konstantin Mishchenko, Ahmed Khaled, and Peter Richtarik · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
Gradient Descent Can Take Exponential Time to Escape Saddle Points, November 2017
Simon S. Du, Chi Jin, Jason D. Lee, Michael I. Jordan, Barnabas Poczos, and Aarti Singh · 2017
Cited alongside, same era.
Topology and Geometry of Half-Rectified Network Optimization
C. Daniel Freeman and Joan Bruna · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
How to Escape Saddle Points Efficiently, March 2017
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Cited alongside, same era.
Exploring Generalization in Deep Learning
Behnam Neyshabur, Srinadh Bhojanapalli, David Mcallester, and Nati Srebro · 2017
Cited alongside, same era.
Henning Petzka and Cristian Sminchisescu · 2020
Later among the works it cites.
Optimization for Deep Learning: An Overview
Ruo-Yu Sun · 2020
Later among the works it cites.
On the interplay between noise and curvature and its effect on optimization and generalization
Valentin Thomas, Fabian Pedregosa, Bart Merriënboer, Pierre-Antoine Manzagol, Yoshua Bengio, and Nicolas Le Roux · 2020
Later among the works it cites.
Spurious Valleys in Two-layer Neural Network Optimization Landscapes
Luca Venturi, Afonso S. Bandeira, and Joan Bruna · 2020
Later among the works it cites.
Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability
Jeremy M. Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet Talwalkar · 2021
Later among the works it cites.
Label Noise SGD Provably Prefers Flat Global Minimizers
Alex Damian, Tengyu Ma, and Jason Lee · 2021
Later among the works it cites.
Stochastic Training is Not Necessary for Generalization
Jonas Geiping, Micah Goldblum, Phillip E. Pope, Michael Moeller, and Tom Goldstein · 2021
Later among the works it cites.
Why Random Reshuffling Beats Stochastic Gradient Descent
Mert Gürbüzbalaban, Asuman Ozdaglar, and Pablo Parrilo · 2021
Later among the works it cites.
On Nonconvex Optimization for Machine Learning: Gradients, Stochasticity, and Saddle Points
Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M. Kakade, and Michael I. Jordan · 2021
Later among the works it cites.
On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)
Zhiyuan Li, Sadhika Malladi, and Sanjeev Arora · 2021
Later among the works it cites.
SGD Implicitly Regularizes Generalization Error
Daniel A. Roberts · 2021
Later among the works it cites.
On the Origin of Implicit Regularization in Stochastic Gradient Descent
Samuel L. Smith, Benoit Dherin, David G. T. Barrett, and Soham De · 2021
Later among the works it cites.
Implicit Regularization in ReLU Networks with the Square Loss
Gal Vardi and Ohad Shamir · 2021
Later among the works it cites.
Understanding Deep Learning (Still) Requires Rethinking Generalization | March 2021 | Communications of the ACM, 2021
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
Adaptive Gradient Methods at the Edge of Stability, July 2022
Jeremy M. Cohen, Behrooz Ghorbani, Shankar Krishnan, Naman Agarwal, Sourabh Medapati, Michal Badura, Daniel Suo, David Cardoze, Zachary Nado, George E. Dahl, and Justin Gilmer · 2022
Later among the works it cites.
What Happens after SGD Reaches Zero Loss? –A Mathematical Framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2022
Later among the works it cites.
Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets, January 2022
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Later among the works it cites.
Embedding Principle of Loss Landscape of Deep Neural Networks, January 2022
Yaoyu Zhang, Zhongwang Zhang, Tao Luo, and Zhi-Qin John Xu · 2022
Later among the works it cites.
On the Implicit Bias of Adam, October 2023
Matias D. Cattaneo, Jason M. Klusowski, and Boris Shigida · 2023
Closest in time.
Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks, June 2023
Feng Chen, Daniel Kunin, Atsushi Yamamura, and Surya Ganguli · 2023
Closest in time.
Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of Stability, April 2023
Alex Damian, Eshaan Nichani, and Jason D. Lee · 2023
Closest in time.
Avrajit Ghosh, He Lyu, Xitong Zhang, and Rongrong Wang · 2023
Closest in time.
Benign Oscillation of Stochastic Gradient Descent with Large Learning Rates, October 2023
Miao Lu, Beining Wu, Xiaodong Yang, and Difan Zou · 2023
Closest in time.
Explicit Regularization in Overparametrized Models via Noise Injection, January 2023
Antonio Orvieto, Anant Raj, Hans Kersting, and Francis Bach · 2023
Closest in time.