Fetching the paper…
Reading the bibliography…
Previous work has shown that DNNs with large depth $L$ and $L_{2}$-regularization are biased towards learning low-dimensional representations of the inputs, which can be interpreted as minimizing a notion of rank $R^{(0)}(f)$ of the learned function $f$, conjectured to be the Bottleneck rank.
The shortest path through many points
Jillian Beardwood, J. H. Halton, and J. M. Hammersley · 1959
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Understanding deep neural networks with rectified linear units
Raman Arora, Amitabh Basu, Poorya Mianjy, and Anirbit Mukherjee · 2018
Earlier work this paper cites.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Earlier work this paper cites.
Relu deep neural networks and linear finite elements
Juncai He, Lin Li, Jinchao Xu, and Chunyue Zheng · 2018
Earlier work this paper cites.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lénaïc Chizat and Francis Bach · 2020
Cited alongside, same era.
The asymptotic spectrum of the hessian of dnn throughout training
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2020
Cited alongside, same era.
Do ideas have shape? plato’s theory of forms as the continuous limit of artificial neural networks
Houman Owhadi · 2020
Cited alongside, same era.
Representation costs of linear neural networks: Analysis and design
Zhen Dai, Mina Karzand, and Nathan Srebro · 2021
Cited alongside, same era.
High-dimensional limit theorems for SGD: Effective dynamics and critical scaling
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2022
Later among the works it cites.
Feature learning in l 2 l_{2} -regularized dnns: Attraction/repulsion and sparsity
Arthur Jacot, Eugene Golikov, Clément Hongler, and Franck Gabriel · 2022
Later among the works it cites.
Training invariances and the low-rank phenomenon: beyond linear networks
Thien Le and Stefanie Jegelka · 2022
Later among the works it cites.
The role of linear layers in nonlinear interpolating networks
Greg Ongie and Rebecca Willett · 2022
Later among the works it cites.
Implicit regularization towards rank minimization in relu networks
Nadav Timor, Gal Vardi, and Ohad Shamir · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Damian, Tengyu Ma, and Jason D Lee · 2021
Cited alongside, same era.
What happens after sgd reaches zero loss?–a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2021
Cited alongside, same era.
Implicit bias of large depth networks: a notion of rank for nonlinear functions
Arthur Jacot · 2023
Closest in time.
Implicit bias of sgd in l 2 l_{2} -regularized linear dnns: One-way jumps from high to low rank, 2023
Zihan Wang and Arthur Jacot · 2023
Closest in time.