Fetching the paper…
Reading the bibliography…
This paper explores the implicit bias of overparameterized neural networks of depth greater than two layers.
A function space view of bounded norm infinite width relu nets: The multivariate case
Ongie, G., Willett, R., Soudry, D., and Srebro, N. (2019) · 1910
Earlier work this paper cites.
Comparing biases for minimal network construction with back-propagation
Hanson, S. and Pratt, L. (1988) · 1988
Earlier work this paper cites.
Maximum-margin matrix factorization
Srebro, N., Rennie, J. D., and Jaakkola, T. S. (2004) · 2004
Earlier work this paper cites.
Implicit regularization in deep learning may not be explainable by norms
Razin, N. and Cohen, N. (2020) · 2005
Earlier work this paper cites.
Computation of matrix norms with applications to robust optimization
Steinberg, D. (2005) · 2005
Earlier work this paper cites.
Neural networks with small weights and depth-separation barriers
Vardi, G. and Shamir, O. (2020) · 2006
Earlier work this paper cites.
Model selection and estimation in regression with grouped variables
Yuan, M. and Lin, Y. (2006) · 2006
Earlier work this paper cites.
Are wider nets better given the same number of parameters?
Golubeva, A., Neyshabur, B., and Gur-Ari, G. (2020) · 2010
Earlier work this paper cites.
Do deep nets really need to be deep?
Ba, L. J. and Caruana, R. (2013) · 2013
Earlier work this paper cites.
Norm-based capacity control in neural networks
Neyshabur, B., Tomioka, R., and Srebro, N. (2015) · 2015
Cited alongside, same era.
Do deep convolutional nets really need to be deep and convolutional?
Urban, G., Geras, K. J., Kahou, S. E., Aslan, O., Wang, S., Caruana, R., Mohamed, A., Philipose, M., and Richardson, M. (2016) · 2016
Cited alongside, same era.
Depth separation for neural networks
Daniely, A. (2017) · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2017) · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., Mcallester, D., and Srebro, N. (2017) · 2017
Cited alongside, same era.
Universal function approximation by deep neural nets with bounded width and relu activations
Hanin, B. (2019) · 2019
Later among the works it cites.
Depth separations in neural networks: what is actually being separated?
Safran, I., Eldan, R., and Shamir, O. (2019) · 2019
Later among the works it cites.
How do infinite width bounded norm networks look in function space?
Savarese, P., Evron, I., Soudry, D., and Srebro, N. (2019) · 2019
Later among the works it cites.
A unified scalable equivalent formulation for Schatten quasi-norms
Shang, F., Liu, Y., Shang, F., Liu, H., Kong, L., and Jiao, L. (2020) · 2020
Later among the works it cites.
Representation costs of linear neural networks: Analysis and design
Dai, Z., Karzand, M., and Srebro, N. (2021) · 2021
Later among the works it cites.
The implicit bias of minima stability: A view from function space
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arora, S., Cohen, N., and Hazan, E. (2018) · 2018
Cited alongside, same era.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B., Bhojanapalli, S., Neyshabur, B., and Srebro, N. (2018) · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Arora, S., Cohen, N., Hu, W., and Luo, Y. (2019) · 2019
Cited alongside, same era.
Mulayoff, R., Michaeli, T., and Soudry, D. (2021) · 2021
Later among the works it cites.
Banach space representer theorems for neural networks and ridge splines
Parhi, R. and Nowak, R. D. (2021) · 2021
Later among the works it cites.
Implicit regularization in tensor factorization
Razin, N., Maman, A., and Cohen, N. (2021) · 2021
Later among the works it cites.