Fetching the paper…
Reading the bibliography…
Understanding the implicit bias of Stochastic Gradient Descent (SGD) is one of the key challenges in deep learning, especially for overparametrized models, where the local minimizers of the loss function $L$ can form a manifold.
On generalization error bounds of noisy gradient methods for non-convex learning
Jian Li, Xuanyuan Luo, and Mingda Qiao · 1902
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 1905
Earlier work this paper cites.
Yuanzhi Li, Colin Wei, and Tengyu Ma · 1907
Earlier work this paper cites.
A topological property of real analytic subsets
Stanislaw Lojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for solving equations and inequalities
Boris T Polyak · 1964
Earlier work this paper cites.
Differentiation of the limit mapping in a dynamical system
K. J. Falconer · 1983
Earlier work this paper cites.
Solutions of a stochastic differential equation forced onto a manifold by a large drift
Gary Shon Katzenberger · 1991
Earlier work this paper cites.
On the probability that a random ± \pm 1-matrix is singular
Jeff Kahn, János Komlós, and Endre Szemerédi · 1995
Earlier work this paper cites.
Differential equations and dynamical systems
Lawrence M. Perko · 2001
Earlier work this paper cites.
Stochastic analysis on manifolds
Elton P Hsu · 2002
Earlier work this paper cites.
Stochastic-process limits: an introduction to stochastic-process limits and their application to queues
Ward Whitt · 2002
Earlier work this paper cites.
Feature selection, l 1 vs. l 2 regularization, and rotational invariance
Andrew Y Ng · 2004
Earlier work this paper cites.
Why are convolutional nets more sample-efficient than fully-connected nets?
Zhiyuan Li, Yi Zhang, and Sanjeev Arora · 2010
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Convergence of stochastic processes
David Pollard · 2012
Earlier work this paper cites.
Minimax-optimal rates for sparse additive models over kernel classes via convex programming
Garvesh Raskutti, Martin J Wainwright, and Bin Yu · 2012
Earlier work this paper cites.
Lectures on Morse homology , volume 29
Augustin Banyaga and David Hurtubise · 2013
Earlier work this paper cites.
Convergence of probability measures
Patrick Billingsley · 2013
Earlier work this paper cites.
Riemannian geometry
Manfredo P Do Carmo · 2013
Earlier work this paper cites.
Brownian motion and stochastic calculus , volume 113
Ioannis Karatzas and Steven Shreve · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Earlier work this paper cites.
Convex recovery of a structured signal from independent random linear measurements
Joel A Tropp · 2015
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Rohith Kuditipudi, Xiang Wang, Holden Lee, Yi Zhang, Zhiyuan Li, Wei Hu, Sanjeev Arora, and Rong Ge · 2019
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Later among the works it cites.
On connected sublevel sets in deep learning
Quynh Nguyen · 2019
Later among the works it cites.
Implicit regularization for optimal sparse recovery
Tomas Vaskevicius, Varun Kanade, and Patrick Rebeschini · 2019
Later among the works it cites.
Yeming Wen, Kevin Luk, Maxime Gazeau, Guodong Zhang, Harris Chan, and Jimmy Ba · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
First-order methods almost always avoid saddle points
Jason D Lee, Ioannis Panageas, Georgios Piliouras, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2017
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and E Weinan · 2017
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2018
Cited alongside, same era.
The loss landscape of overparameterized neural networks
Yaim Cooper · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
Greg Yang · 2019
Later among the works it cites.
Peng Zhao, Yun Yang, and Qiao-Chu He · 2019
Later among the works it cites.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Later among the works it cites.
Stochastic gradient and langevin processes
Xiang Cheng, Dong Yin, Peter Bartlett, and Michael Jordan · 2020
Later among the works it cites.
The critical locus of overparameterized neural networks
Y Cooper · 2020
Later among the works it cites.
Understanding implicit regularization in over-parameterized nonlinear statistical model
Jianqing Fan, Zhuoran Yang, and Mengxin Yu · 2020
Later among the works it cites.
Convergence rates for the stochastic gradient descent method for non-convex objective functions
Benjamin Fehrman, Benjamin Gess, and Arnulf Jentzen · 2020
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z HaoChen, Colin Wei, Jason D Lee, and Tengyu Ma · 2020
Later among the works it cites.
On learning rates and schr \ \backslash ” odinger operators
Bin Shi, Weijie J Su, and Michael I Jordan · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
On the noisy gradient descent that generalizes as sgd
Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, and Zhanxing Zhu · 2020
Later among the works it cites.
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
On the implicit bias of initialization shape: Beyond infinitesimal mirror descent
Shahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake Woodworth, Nathan Srebro, Amir Globerson, and Daniel Soudry · 2021
Closest in time.
Label noise sgd provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason Lee · 2021
Closest in time.
On the validity of modeling sgd with stochastic differential equations (sdes)
Zhiyuan Li, Sadhika Malladi, and Sanjeev Arora · 2021
Closest in time.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Closest in time.
Stochastic gradient descent with noise of machine learning type. part ii: Continuous time analysis
Stephan Wojtowytsch · 2021
Closest in time.