Fetching the paper…
Reading the bibliography…
We show that the representation cost of fully connected neural networks with homogeneous nonlinearities - which describes the implicit bias in function space of networks with $L_2$-regularization or with losses such as the cross-entropy - converges as the depth of the network goes to infinity to a notion of rank over nonlinear functions.
The shortest path through many points
Jillian Beardwood, J. H. Halton, and J. M. Hammersley · 1959
Earlier work this paper cites.
Exact matrix completion via convex optimization
Emmanuel J Candès and Benjamin Recht · 2009
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Understanding deep neural networks with rectified linear units
Raman Arora, Amitabh Basu, Poorya Mianjy, and Anirbit Mukherjee · 2018
Earlier work this paper cites.
On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport
Lénaïc Chizat and Francis Bach · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Earlier work this paper cites.
Relu deep neural networks and linear finite elements
Juncai He, Lin Li, Jinchao Xu, and Chunyue Zheng · 2018
Earlier work this paper cites.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2018
Cited alongside, same era.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
Grant Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 2019
Cited alongside, same era.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Edward Moroshko, Blake E Woodworth, Suriya Gunasekar, Jason D Lee, Nati Srebro, and Daniel Soudry · 2020
Later among the works it cites.
A function space view of bounded norm infinite width relu nets: The multivariate case
Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
Representation costs of linear neural networks: Analysis and design
Zhen Dai, Mina Karzand, and Nathan Srebro · 2021
Later among the works it cites.
The implicit bias of minima stability: A view from function space
Rotem Mulayoff, Tomer Michaeli, and Daniel Soudry · 2021
Later among the works it cites.
Rank diminishing in deep neural networks
Ruili Feng, Kecheng Zheng, Yukun Huang, Deli Zhao, Michael Jordan, and Zheng-Jun Zha · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lénaïc Chizat and Francis Bach · 2020
Cited alongside, same era.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Cited alongside, same era.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2020
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro
Cited in the paper.
Saddle-to-saddle dynamics in deep linear networks: Small initialization training, symmetry, and sparsity, 2022a
Arthur Jacot, François Ged, Berfin Şimşek, Clément Hongler, and Franck Gabriel
Cited in the paper.
Feature learning in l 2 l_{2} -regularized dnns: Attraction/repulsion and sparsity, 2022b
Arthur Jacot, Eugene Golikov, Clément Hongler, and Franck Gabriel
Cited in the paper.
Training invariances and the low-rank phenomenon: beyond linear networks
Thien Le and Stefanie Jegelka · 2022
Closest in time.
The role of linear layers in nonlinear interpolating networks
Greg Ongie and Rebecca Willett · 2022
Closest in time.
Implicit regularization towards rank minimization in relu networks
Nadav Timor, Gal Vardi, and Ohad Shamir · 2022
Closest in time.