Fetching the paper…
Reading the bibliography…
This paper studies the problem of training a two-layer ReLU network for binary classification using gradient flow with small initialization.
Ordinary Differential Equations
W. T. Reid · 1971
Earlier work this paper cites.
The existence of solutions of generalized differential equations
A. F. Filippov · 1971
Earlier work this paper cites.
A generalization of carathéodory’s existence theorem for ordinary differential equations
Jan Persson · 1975
Earlier work this paper cites.
Extension and separation of vector valued functions
Zafer Ercan · 1997
Earlier work this paper cites.
An Introduction to Hybrid Dynamical Systems
A.J. van der Schaft and J.M. Schumacher · 2000
Earlier work this paper cites.
Characterizations of łojasiewicz inequalities: subgradient flows, talweg, convexity
Jérôme Bolte, Aris Daniilidis, Olivier Ley, and Laurent Mazet · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural network
Andrew M Saxe, James L Mcclelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2017
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Earlier work this paper cites.
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Cited alongside, same era.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Cited alongside, same era.
Learning relu networks on linearly separable data: Algorithm, optimality, and generalization
The inductive bias of relu networks on orthogonally separable data
Mary Phuong and Christoph H Lampert · 2021
Later among the works it cites.
Implicit regularization in hierarchical tensor factorization and deep convolutional neural networks
Noam Razin, Asaf Maman, and Nadav Cohen · 2022
Later among the works it cites.
Gradient flow dynamics of shallow relu networks for square loss and orthogonal inputs
Etienne Boursier, Loucas Pullaud-Vivien, and Nicolas Flammarion · 2022
Later among the works it cites.
Training invariances and the low-rank phenomenon: beyond linear networks
Thien Le and Stefanie Jegelka · 2022
Later among the works it cites.
The convex geometry of backpropagation: Neural network gradient flows converge to extreme points of the dual convex program
Yifei Wang and Mert Pilanci · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gang Wang, Georgios B. Giannakis, and Jie Chen · 2019
Cited alongside, same era.
A unifying view on implicit bias in training linear neural networks
Chulhee Yun, Shankar Krishnan, and Hossein Mobahi · 2020
Cited alongside, same era.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Cited alongside, same era.
Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction
Dominik Stöger and Mahdi Soltanolkotabi · 2021
Cited alongside, same era.
On the explicit role of initialization on the convergence and implicit bias of overparametrized linear networks
Hancheng Min, Salma Tarmoun, René Vidal, and Enrique Mallada · 2021
Cited alongside, same era.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2021
Cited alongside, same era.
Gradient descent on two-layer nets: Margin maximization and simplicity bias
Kaifeng Lyu, Zhiyuan Li, Runzhe Wang, and Sanjeev Arora · 2021
Cited alongside, same era.
Mingze Wang and Chao Ma · 2022
Later among the works it cites.
Implicit bias in leaky relu networks trained on high-dimensional data
Spencer Frei, Gal Vardi, Peter Bartlett, Nathan Srebro, and Wei Hu · 2022
Later among the works it cites.
On the spectral bias of two-layer linear networks
Aditya Vardhan Varre, Maria-Luiza Vladarean, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2023
Closest in time.
The law of parsimony in gradient descent for learning deep linear networks, 2023
Can Yaras, Peng Wang, Wei Hu, Zhihui Zhu, Laura Balzano, and Qing Qu · 2023
Closest in time.
Mahdi Soltanolkotabi, Dominik Stöger, and Changzhi Xie · 2023
Closest in time.
Understanding multi-phase optimization dynamics and rich nonlinear behaviors of relu networks
Mingze Wang and Chao Ma · 2023
Closest in time.
Implicit bias of gradient descent for two-layer reLU and leaky reLU networks on nearly-orthogonal data
Yiwen Kou, Zixiang Chen, and Quanquan Gu · 2023
Closest in time.