Fetching the paper…
Reading the bibliography…
The implicit bias induced by the training of neural networks has become a topic of rigorous study.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Kunihiko Fukushima · 1980
Earlier work this paper cites.
Optimization and Nonsmooth Analysis
F.H. Clarke · 1983
Earlier work this paper cites.
The Loss Surfaces of Multilayer Networks
Anna Choromanska, MIkael Henaff, Michael Mathieu, Gerard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Russ R Salakhutdinov, and Nati Srebro · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Geometry of optimization and implicit regularization in deep learning, 2017
Behnam Neyshabur, Ryota Tomioka, Ruslan Salakhutdinov, and Nathan Srebro · 2017
Earlier work this paper cites.
Operator norm inequalities between tensor unfoldings on the partition lattice, May 2017
Miaoyan Wang, Khanh Dao Duc, Jonathan Fischer, and Yun S. Song · 2017
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced, 2018
Simon S. Du, Wei Hu, and Jason D. Lee · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Cited alongside, same era.
Critical points of linear neural networks: Analytical forms and landscape properties
Yi Zhou and Yingbin Liang · 2018
Cited alongside, same era.
Deep relu networks have surprisingly few activation patterns
Adversarial risk bounds via function transformation, 2019
Justin Khim and Po-Ling Loh · 2019
Later among the works it cites.
A mathematical model for automatic differentiation in machine learning
Jérôme Bolte and Edouard Pauwels · 2020
Later among the works it cites.
Stochastic subgradient method converges on tame functions
Damek Davis, Dmitriy Drusvyatskiy, Sham Kakade, and Jason D. Lee · 2020
Later among the works it cites.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Later among the works it cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Later among the works it cites.
On alignment in deep linear neural networks, 2020
Adityanarayanan Radhakrishnan, Eshaan Nichani, Daniel Bernstein, and Caroline Uhler · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Boris Hanin and David Rolnick · 2019
Cited alongside, same era.
Absolute continuity and the fundamental theorem of calculus, 2019
Christopher Heil · 2019
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
Revealing the structure of deep neural networks via convex duality
Tolga Ergen and Mert Pilanci · 2021
Later among the works it cites.
The low-rank simplicity bias in deep networks
Minyoung Huh, Hossein Mobahi, Richard Zhang, Pulkit Agrawal, and Phillip Isola · 2021
Later among the works it cites.