Fetching the paper…
Reading the bibliography…
We study the dynamics and implicit bias of gradient flow (GF) on univariate ReLU neural networks with a single hidden layer in a binary classification setting.
Tables for computing bivariate normal probabilities
D. B. Owen · 1956
Earlier work this paper cites.
Gradient methods for minimizing functionals
B. T. Polyak · 1963
Earlier work this paper cites.
Nonsmooth analysis and control theory , volume 178
F. H. Clarke, Y. S. Ledyaev, R. J. Stern, and P. R. Wolenski · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Approximate kkt points and a proximity measure for termination
J. Dutta, K. Deb, R. Tulshyan, and R. Arora · 2013
Earlier work this paper cites.
Gradient descent optimizes over-parameterized deep relu networks
D. Zou, Y. Cao, D. Zhou, and Q. Gu · 2013
Earlier work this paper cites.
Learning polynomials with neural networks
A. Andoni, R. Panigrahy, G. Valiant, and L. Zhang · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Sgd learns the conjugate kernel class of the network
A. Daniely · 2017
Earlier work this paper cites.
Implicit regularization in deep learning
B. Neyshabur · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro · 2017
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Earlier work this paper cites.
Gradient descent quantizes relu network features
H. Maennel, O. Bousquet, and S. Gelly · 2018
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
I. Safran and O. Shamir · 2018
Cited alongside, same era.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. D. Lee · 2018
Cited alongside, same era.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
S. Du, J. Lee, H. Li, L. Wang, and X. Zhai · 2019
Cited alongside, same era.
Neural tangent kernels, transportation mappings, and universal approximation
The inductive bias of relu networks on orthogonally separable data
M. Phuong and C. H. Lampert · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
N. Razin and N. Cohen · 2020
Later among the works it cites.
J. Sahs, R. Pyle, A. Damaraju, J. O. Caro, O. Tavaslioglu, A. Lu, and A. Patel · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
B. Woodworth, S. Gunasekar, J. D. Lee, E. Moroshko, P. Savarese, I. Golan, D. Soudry, and N. Srebro · 2020
Later among the works it cites.
Analytic study of families of spurious minima in two-layer relu neural networks: A tale of symmetry ii
Y. Arjevani and M. Field · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Ji, M. Telgarsky, and R. Xian · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
K. Lyu and J. Li · 2019
Cited alongside, same era.
A function space view of bounded norm infinite width relu nets: The multivariate case
G. Ongie, R. Willett, D. Soudry, and N. Srebro · 2019
Cited alongside, same era.
How do infinite width bounded norm networks look in function space?
P. Savarese, I. Evron, D. Soudry, and N. Srebro · 2019
Cited alongside, same era.
Gradient dynamics of shallow univariate relu networks
F. Williams, M. Trager, D. Panozzo, C. Silva, D. Zorin, and J. Bruna · 2019
Cited alongside, same era.
Analytic characterization of the hessian in shallow relu models: A tale of symmetry
Y. Arjevani and M. Field · 2020
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
G. Blanc, N. Gupta, G. Valiant, and P. Valiant · 2020
Cited alongside, same era.
NIST Digital Library of Mathematical Functions
DLMF · 2021
Later among the works it cites.
Continuous vs. discrete optimization of deep neural networks
O. Elkabetz and N. Cohen · 2021
Later among the works it cites.
Path length bounds for gradient descent and flow
C. Gupta, S. Balakrishnan, and A. Ramdas · 2021
Later among the works it cites.
B. Hanin · 2021
Later among the works it cites.
The effects of mild over-parameterization on the optimization landscape of shallow relu neural networks
I. M. Safran, G. Yehudai, and O. Shamir · 2021
Later among the works it cites.
Implicit regularization in relu networks with the square loss
G. Vardi and O. Shamir · 2021
Later among the works it cites.
On margin maximization in linear and relu networks
G. Vardi, O. Shamir, and N. Srebro · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2021
Later among the works it cites.
A local convergence theory for mildly over-parameterized two-layer neural network
M. Zhou, R. Ge, and C. Jin · 2021
Later among the works it cites.
Implicit regularization towards rank minimization in relu networks
N. Timor, G. Vardi, and O. Shamir · 2022
Closest in time.
On the implicit bias in deep-learning algorithms
G. Vardi · 2022
Closest in time.