Fetching the paper…
Reading the bibliography…
Gradient-based deep-learning algorithms exhibit remarkable performance in practice, but it is not well-understood why they are able to generalize despite having more parameters than training examples.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2013
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Earlier work this paper cites.
Norm-based capacity control in neural networks
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
P. L. Bartlett, D. J. Foster, and M. J. Telgarsky · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
S. Jastrzębski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Earlier work this paper cites.
Stochastic gradient/mirror descent: Minimax optimality and implicit regularization
N. Azizan and B. Hassibi · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
S. S. Du, W. Hu, and J. D. Lee · 2018
Earlier work this paper cites.
Size-independent sample complexity of neural networks
N. Golowich, A. Rakhlin, and O. Shamir · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Y. Li, T. Ma, and H. Zhang · 2018
Earlier work this paper cites.
Step size matters in deep learning
K. Nar and S. Sastry · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
L. Wu, C. Ma, et al · 2018
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
S. Arora, N. Cohen, W. Hu, and Y. Luo · 2019
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
M. Belkin, D. Hsu, S. Ma, and S. Mandal · 2019
Earlier work this paper cites.
On the inductive bias of neural tangent kernels
A. Bietti and J. Mairal · 2019
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
K. Lyu and J. Li · 2019
Earlier work this paper cites.
A function space view of bounded norm infinite width relu nets: The multivariate case
G. Ongie, R. Willett, D. Soudry, and N. Srebro · 2019
Earlier work this paper cites.
How do infinite width bounded norm networks look in function space?
P. Savarese, I. Evron, D. Soudry, and N. Srebro · 2019
Earlier work this paper cites.
Gradient dynamics of shallow univariate relu networks
F. Williams, M. Trager, D. Panozzo, C. Silva, D. Zorin, and J. Bruna · 2019
Earlier work this paper cites.
Implicit gradient regularization
D. G. Barrett and B. Dherin · 2020
Earlier work this paper cites.
On implicit regularization: Morse functions and applications to matrix factorization
M. A. Belabbas · 2020
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
G. Blanc, N. Gupta, G. Valiant, and P. Valiant · 2020
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
L. Chizat and F. Bach · 2020
Cited alongside, same era.
Directional convergence and alignment in deep learning
Z. Ji and M. Telgarsky · 2020
Cited alongside, same era.
Gradient descent follows the regularization path for general losses
Z. Ji, M. Dudík, R. E. Schapire, and M. Telgarsky · 2020
Cited alongside, same era.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
S. Pesme, L. Pillaud-Vivien, and N. Flammarion · 2021
Later among the works it cites.
Implicit regularization in tensor factorization
N. Razin, A. Maman, and N. Cohen · 2021
Later among the works it cites.
Towards understanding learning in neural networks with linear teachers
R. Sarussi, A. Brutzkus, and A. Globerson · 2021
Later among the works it cites.
A theoretical analysis of fine-tuning with linear teachers
G. Shachaf, A. Brutzkus, and A. Globerson · 2021
Later among the works it cites.
Gradient methods never overfit on separable data
O. Shamir · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
S. L. Smith, B. Dherin, D. G. Barrett, and S. De · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Li, Y. Luo, and K. Lyu · 2020
Cited alongside, same era.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
E. Moroshko, B. E. Woodworth, S. Gunasekar, J. D. Lee, N. Srebro, and D. Soudry · 2020
Cited alongside, same era.
Unique properties of flat minima in deep networks
R. Mulayoff and T. Michaeli · 2020
Cited alongside, same era.
Neural networks are convex regularizers: Exact polynomial-time convex optimization formulations for two-layer networks
M. Pilanci and T. Ergen · 2020
Cited alongside, same era.
Implicit regularization in deep learning may not be explainable by norms
N. Razin and N. Cohen · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
B. Woodworth, S. Gunasekar, J. D. Lee, E. Moroshko, P. Savarese, I. Golan, D. Soudry, and N. Srebro · 2020
Cited alongside, same era.
A unifying view on implicit bias in training linear neural networks
C. Yun, S. Krishnan, and H. Mobahi · 2020
Cited alongside, same era.
Later among the works it cites.
Implicit regularization in relu networks with the square loss
G. Vardi and O. Shamir · 2021
Later among the works it cites.
On margin maximization in linear and relu networks
G. Vardi, O. Shamir, and N. Srebro · 2021
Later among the works it cites.
Large learning rate tames homogeneity: Convergence and balancing effect
Y. Wang, M. Chen, T. Zhao, and M. Tao · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2021
Later among the works it cites.
Understanding the unstable convergence of gradient descent
K. Ahn, J. Zhang, and S. Sra · 2022
Closest in time.
Understanding gradient descent on edge of stability in deep learning
S. Arora, Z. Li, and A. Panigrahi · 2022
Closest in time.
Gradient flow dynamics of shallow relu networks for square loss and orthogonal inputs
E. Boursier, L. Pillaud-Vivien, and N. Flammarion · 2022
Closest in time.
On gradient descent convergence beyond the edge of stability
L. Chen and J. Bruna · 2022
Closest in time.
Implicit bias in leaky relu networks trained on high-dimensional data
S. Frei, G. Vardi, P. L. Bartlett, N. Srebro, and W. Hu · 2022
Closest in time.
Reconstructing training data from trained neural networks
N. Haim, G. Vardi, G. Yehudai, O. Shamir, and M. Irani · 2022
Closest in time.
Inductive bias of multi-channel linear convolutional networks with bounded weight norm
M. Jagadeesan, I. Razenshteyn, and S. Gunasekar · 2022
Closest in time.
Understanding the generalization benefit of normalization layers: Sharpness reduction
K. Lyu, Z. Li, and S. Arora · 2022
Closest in time.
Beyond the quadratic approximation: the multiscale structure of neural network loss landscapes
C. Ma, L. Wu, and L. Ying · 2022
Closest in time.
Implicit bias of the step size in linear diagonal neural networks
M. S. Nacson, K. Ravichandran, N. Srebro, and D. Soudry · 2022
Closest in time.
Label noise (stochastic) gradient descent implicitly solves the lasso for quadratic parametrisation
L. Pillaud-Vivien, J. Reygner, and N. Flammarion · 2022
Closest in time.
Implicit regularization in hierarchical tensor factorization and deep convolutional neural networks
N. Razin, A. Maman, and N. Cohen · 2022
Closest in time.
I. Safran, G. Vardi, and J. D. Lee · 2022
Closest in time.
Implicit regularization towards rank minimization in relu networks
N. Timor, G. Vardi, and O. Shamir · 2022
Closest in time.
When does sgd favor flat minima? a quantitative characterization via linear stability
L. Wu, M. Wang, and W. Su · 2022
Closest in time.