Fetching the paper…
Reading the bibliography…
Neural networks often operate in the overparameterized regime, in which there are far more parameters than training samples, allowing the training data to be fit perfectly.
Comparing biases for minimal network construction with back-propagation
Hanson, S. and Pratt, L. (1988) · 1988
Earlier work this paper cites.
Exploring regression structure using nonparametric functional estimation
Samarov, A. M. (1993) · 1993
Earlier work this paper cites.
For valid generalization the size of the weights is more important than the size of the network
Bartlett, P. L. (1997) · 1997
Earlier work this paper cites.
Some inequalities for singular values of matrix products
Wang, B.-Y. and Xi, B.-Y. (1997) · 1997
Earlier work this paper cites.
High-dimensional data analysis: The curses and blessings of dimensionality
Donoho, D. L. (2000) · 2000
Earlier work this paper cites.
Structure adaptive approach for dimension reduction
Hristache, M., Juditsky, A., Polzehl, J., and Spokoiny, V. (2001) · 2001
Earlier work this paper cites.
Maximum-margin matrix factorization
Srebro, N., Rennie, J., and Jaakkola, T. (2004) · 2004
Earlier work this paper cites.
Fourier methods for estimating the central subspace and the central mean subspace in regression
Zhu, Y. and Zeng, P. (2006) · 2006
Earlier work this paper cites.
A multiple-index model and dimension reduction
Xia, Y. (2008) · 2008
Earlier work this paper cites.
Successive direction extraction for estimating the central subspace in a multiple-index regression
Yin, X., Li, B., and Cook, R. D. (2008) · 2008
Earlier work this paper cites.
Partial differential equations
Evans, L. C. (2010) · 2010
Earlier work this paper cites.
Learning gradients: predictive models that infer geometry and statistical dependence
Wu, Q., Guinney, J., Maggioni, M., and Mukherjee, S. (2010) · 2010
Earlier work this paper cites.
Efficient learning of generalized linear and single index models with isotonic regression
Kakade, S. M., Kanade, V., Shamir, O., and Kalai, A. (2011) · 2011
Earlier work this paper cites.
Capturing ridge functions in high dimensions from point queries
Cohen, A., Daubechies, I., DeVore, R., Kerkyacharian, G., and Picard, D. (2012) · 2012
Earlier work this paper cites.
Low-rank matrix recovery via efficient schatten p-norm minimization
Nie, F., Huang, H., and Ding, C. (2012) · 2012
Earlier work this paper cites.
Do deep nets really need to be deep?
Ba, J. and Caruana, R. (2014) · 2014
Earlier work this paper cites.
Mathematics of sparsity (and a few other things)
Candès, E. J. (2014) · 2014
Earlier work this paper cites.
Active subspace methods in theory and practice: Applications to kriging surfaces
Constantine, P. G., Dow, E., and Wang, Q. (2014) · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N. (2014) · 2014
Earlier work this paper cites.
A consistent estimator of the expected gradient outerproduct
Trivedi, S., Wang, J., Kpotufe, S., and Shakhnarovich, G. (2014) · 2014
Earlier work this paper cites.
Active subspaces: Emerging ideas for dimension reduction in parameter studies
Constantine, P. G. (2015) · 2015
Earlier work this paper cites.
Matrix completion under monotonic single index models
Ganti, R. S., Balzano, L., and Willett, R. (2015) · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Neyshabur, B., Tomioka, R., and Srebro, N. (2015) · 2015
Earlier work this paper cites.
Scalable algorithms for tractable schatten quasi-norm minimization
Shang, F., Liu, Y., and Cheng, J. (2016) · 2016
Earlier work this paper cites.
Do deep convolutional nets really need to be deep and convolutional?
Urban, G., Geras, K. J., Kahou, S. E., Aslan, O., Wang, S., Mohamed, A., Philipose, M., Richardson, M., and Caruana, R. (2016) · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Bach, F. (2017) · 2017
Earlier work this paper cites.
On learning high dimensional structured single index models
Ganti, R., Rao, N., Balzano, L., Willett, R., and Nowak, R. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E. (2018) · 2018
Cited alongside, same era.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B., Bhojanapalli, S., Neyshabur, B., and Srebro, N. (2018) · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Arora, S., Cohen, N., Hu, W., and Luo, Y. (2019) · 2019
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ji, Z. and Telgarsky, M. (2019) · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2019) · 2019
Cited alongside, same era.
A priori estimates of the population risk for two-layer neural networks
Ma, C., Wu, L., et al. (2019) · 2019
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2022) · 2022
Later among the works it cites.
The low-rank simplicity bias in deep networks
Huh, M., Mobahi, H., Zhang, R., Cheung, B., Agrawal, P., and Isola, P. (2022) · 2022
Later among the works it cites.
Uniform approximation rates and metric entropy of shallow neural networks
Ma, L., Siegel, J. W., and Xu, J. (2022) · 2022
Later among the works it cites.
Neural networks efficiently learn low-dimensional representations with SGD
Mousavi-Hosseini, A., Park, S., Girotti, M., Mitliagkas, I., and Erdogdu, M. A. (2022) · 2022
Later among the works it cites.
The implicit bias of minima stability in multivariate shallow relu networks
Nacson, M. S., Mulayoff, R., Ongie, G., Michaeli, T., and Soudry, D. (2022) · 2022
Later among the works it cites.
Parallel deep neural networks have zero duality gap
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lexicographic and depth-sensitive margins in homogeneous and non-homogeneous deep models
Nacson, M. S., Gunasekar, S., Lee, J. D., Srebro, N., and Soudry, D. (2019) · 2019
Cited alongside, same era.
How do infinite width bounded norm networks look in function space?
Savarese, P., Evron, I., Soudry, D., and Srebro, N. (2019) · 2019
Cited alongside, same era.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Wei, C., Lee, J. D., Liu, Q., and Ma, T. (2019) · 2019
Cited alongside, same era.
Implicit convex regularizers of CNN architectures: Convex optimization of two-and three-layer networks in polynomial time
Ergen, T. and Pilanci, M. (2020) · 2020
Cited alongside, same era.
Initialization and regularization of factorized neural layers
Khodak, M., Tenenholtz, N. A., Mackey, L., and Fusi, N. (2020) · 2020
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J. (2020) · 2020
Cited alongside, same era.
Wang, Y., Ergen, T., and Pilanci, M. (2022) · 2022
Later among the works it cites.
Intrinsic dimensionality and generalization properties of the r-norm inductive bias
Ardeshir, N., Hsu, D. J., and Sanford, C. H. (2023) · 2023
Closest in time.
Understanding neural networks with reproducing kernel banach spaces
Bartolucci, F., De Vito, E., Rosasco, L., and Vigogna, S. (2023) · 2023
Closest in time.
Penalising the biases in norm regularisation enforces sparsity
Boursier, E. and Flammarion, N. (2023) · 2023
Closest in time.
(s)gd over diagonal linear networks: Implicit bias, large stepsizes and edge of stability
Even, M., Pesme, S., Gunasekar, S., and Flammarion, N. (2023) · 2023
Closest in time.
Implicit bias of large depth networks: A notion of rank for nonlinear functions
Jacot, A. (2023) · 2023
Closest in time.
Vector-valued variation spaces and width bounds for DNNs: Insights on weight decay regularization
Shenouda, J., Parhi, R., Lee, K., and Nowak, R. D. (2023) · 2023
Closest in time.
Characterization of the variation spaces corresponding to shallow neural networks
Siegel, J. W. and Xu, J. (2023) · 2023
Closest in time.
Ridges, neural networks, and the radon transform
Unser, M. (2023) · 2023
Closest in time.
Efficient estimation of the central mean subspace via smoothed gradient outer products
Yuan, G., Xu, M., Kpotufe, S., and Hsu, D. (2023) · 2023
Closest in time.
Autoencoders for discovering manifold dimension and coordinates in data from complex dynamical systems
Zeng, K., Perez De Jesus, C. E., Fox, A. J., and Graham, M. D. (2023) · 2023
Closest in time.
Average gradient outer product as a mechanism for deep neural collapse
Beaglehole, D., Súkeník, P., Mondelli, M., and Belkin, M. (2024) · 2024
Closest in time.
Neural hilbert ladders: Multi-layer neural networks in function space
Chen, Z. (2024) · 2024
Closest in time.
Path regularization: A convexity and sparsity inducing regularization for parallel relu networks
Ergen, T. and Pilanci, M. (2024) · 2024
Closest in time.
Agnostically learning single-index models using omnipredictors
Gollakota, A., Gopalan, P., Klivans, A., and Stavropoulos, K. (2024) · 2024
Closest in time.
Bottleneck structure in learned features: Low-dimension vs regularity tradeoff
Jacot, A. (2024) · 2024
Closest in time.
Learning functions varying along a central subspace
Liu, H. and Liao, W. (2024) · 2024
Closest in time.
Sharp bounds on the approximation rates, metric entropy, and n-widths of shallow neural networks
Siegel, J. W. and Xu, J. (2024) · 2024
Closest in time.
Implicit bias of sgd in l2-regularized linear dnns: One-way jumps from high to low rank
Wang, Z. and Jacot, A. (2024) · 2024
Closest in time.
Optimal bump functions for shallow ReLU networks: Weight decay, depth separation, curse of dimensionality
Wojtowytsch, S. (2024) · 2024
Closest in time.
Compressible dynamics in deep overparameterized low-rank learning &; adaptation
Yaras, C., Wang, P., Balzano, L., and Qu, Q. (2024) · 2024
Closest in time.
Zeger, E., Wang, Y., Mishkin, A., Ergen, T., Candès, E., and Pilanci, M. (2024) · 2024
Closest in time.
Principal angles between subspaces in an A-based scalar product: algorithms and perturbation estimates
Knyazev, A. V. and Argentati, M. E. (2002) · 2040
Closest in time.