Fetching the paper…
Reading the bibliography…
In recent years neural networks have achieved impressive results on many technological and scientific tasks.
On the generalized distance in statistics
P. C. Mahalanobis · 1936
Earlier work this paper cites.
The statistical utilization of multiple measurements
R. A. Fisher · 1938
Earlier work this paper cites.
Theory of reproducing kernels
N. Aronszajn · 1950
Earlier work this paper cites.
Low-dimensional procedure for the characterization of human faces
L. Sirovich and M. Kirby · 1987
Earlier work this paper cites.
Investigating smooth multiple regression by the method of average derivatives
W. Härdle and T. M. Stoker · 1989
Earlier work this paper cites.
Networks for approximation and learning
T. Poggio and F. Girosi · 1990
Earlier work this paper cites.
Spline models for observational data
G. Wahba · 1990
Earlier work this paper cites.
Face recognition using eigenfaces
M. A. Turk and A. P. Pentland · 1991
Earlier work this paper cites.
On principal hessian directions for data visualization and dimension reduction: Another application of stein’s lemma
K.-C. Li · 1992
Earlier work this paper cites.
The expectation-maximization algorithm
T. K. Moon · 1996
Earlier work this paper cites.
Bayesian Learning for Neural Networks
R. Neal · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
Eigenfaces vs. Fisherfaces: Recognition using class specific linear projection
P. N. Belhumeur, J. P. Hespanha, and D. J. Kriegman · 1997
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Y. Freund and R. E. Schapire · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Nonlinear dimensionality reduction by locally linear embedding
S. T. Roweis and L. K. Saul · 2000
Earlier work this paper cites.
On kernel-target alignment
N. Cristianini, J. Shawe-Taylor, A. Elisseeff, and J. Kandola · 2001
Earlier work this paper cites.
Structure adaptive approach for dimension reduction
M. Hristache, A. Juditsky, J. Polzehl, and V. Spokoiny · 2001
Earlier work this paper cites.
Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond
B. Schölkopf and A. J. Smola · 2002
Earlier work this paper cites.
An adaptive estimation of dimension reduction space
Y. Xia, H. Tong, W. K. Li, and L.-X. Zhu · 2002
Earlier work this paper cites.
Laplacian eigenmaps for dimensionality reduction and data representation
M. Belkin and P. Niyogi · 2003
Earlier work this paper cites.
Estimation of gradients and coordinate covariation in classification
S. Mukherjee and Q. Wu · 2006
Earlier work this paper cites.
Learning coordinate covariances via gradients
S. Mukherjee, D.-X. Zhou, and J. Shawe-Taylor · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
An analysis of single layer networks in unsupervised feature learning
A. Coates, H. Lee, and A. Y. Ng · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng · 2011
Earlier work this paper cites.
Scikit-learn: Machine Learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
The numpy array: a structure for efficient numerical computation
S. Van Der Walt, S. C. Colbert, and G. Varoquaux · 2011
Earlier work this paper cites.
Algorithms for learning kernels based on centered alignment
C. Cortes, M. Mohri, and A. Rostamizadeh · 2012
Earlier work this paper cites.
Do we need hundreds of classifiers to solve real world classification problems?
M. Fernández-Delgado, E. Cernadas, S. Barro, and D. Amorim · 2014
Earlier work this paper cites.
A consistent estimator of the expected gradient outerproduct
S. Trivedi, J. Wang, S. Kpotufe, and G. Shakhnarovich · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
Confidence intervals for low dimensional parameters in high dimensional linear models
C.-H. Zhang and S. S. Zhang · 2014
Earlier work this paper cites.
Metric learning
A. Bellet, A. Habrard, and M. Sebban · 2015
Cited alongside, same era.
Active subspaces: Emerging ideas for dimension reduction in parameter studies
P. G. Constantine · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Deep learning face attributes in the wild
Z. Liu, P. Luo, X. Wang, and X. Tang · 2015
Cited alongside, same era.
An overview of kernel alignment and its applications
T. Wang, D. Zhao, and S. Tian · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolution network, 2015
B. Xu, N. Wang, T. Chen, and M. Li · 2015
Cited alongside, same era.
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel
S. Fort, G. K. Dziugaite, M. Paul, S. Kharaghani, D. M. Roy, and S. Ganguli · 2020
Later among the works it cites.
Finite depth and width corrections to the Neural Tangent Kernel
B. Hanin and M. Nica · 2020
Later among the works it cites.
What shapes feature representations? exploring datasets, architectures, and training
K. Hermann and A. Lampinen · 2020
Later among the works it cites.
Dynamics of deep neural networks and neural tangent hierarchy
J. Huang and H.-T. Yau · 2020
Later among the works it cites.
Kernel alignment risk estimator: Risk prediction from training data
A. Jacot, B. Simsek, F. Spadaro, C. Hongler, and F. Gabriel · 2020
Later among the works it cites.
The large learning rate phase of deep learning: the catapult mechanism
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
XGBoost: A scalable tree boosting system
T. Chen and C. Guestrin · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Gradients weights improve regression and classification
S. Kpotufe, A. Boularias, T. Schultz, and K. Kim · 2016
Cited alongside, same era.
S. Schoenholz, J. Gilmer, S. Ganguli, and J. Sohl-Dickstein · 2016
Cited alongside, same era.
Learning kernels with random features
A. Sinha and J. C. Duchi · 2016
Cited alongside, same era.
Diving into the shallows: a computational perspective on large-scale shallow learning
S. Ma and M. Belkin · 2017
Cited alongside, same era.
A. Lewkowycz, Y. Bahri, E. Dyer, J. Sohl-Dickstein, and G. Gur-Ari · 2020
Later among the works it cites.
On the linearity of large non-linear models: when and why the tangent kernel is constant
C. Liu, L. Zhu, and M. Belkin · 2020
Later among the works it cites.
The pitfalls of simplicity bias in neural networks
H. Shah, K. Tamuly, A. Raghunathan, P. Jain, and P. Netrapalli · 2020
Later among the works it cites.
Linearized two-layers neural networks in high dimension
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2021
Later among the works it cites.
The low-rank simplicity bias in deep networks
M. Huh, H. Mobahi, R. Zhang, B. Cheung, P. Agrawal, and P. Isola · 2021
Later among the works it cites.
De-biased sparse PCA: Inference for eigenstructure of large covariance matrices
J. Janková and S. van de Geer · 2021
Later among the works it cites.
Properties of the after kernel
P. M. Long · 2021
Later among the works it cites.
Gradient starvation: A learning proclivity in neural networks
M. Pezeshki, O. Kaba, Y. Bengio, A. C. Courville, D. Precup, and G. Lajoie · 2021
Later among the works it cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Later among the works it cites.
Saint: Improved neural networks for tabular data via row attention and contrastive pre-training
G. Somepalli, M. Goldblum, A. Schwarzschild, C. B. Bruss, and T. Goldstein · 2021
Later among the works it cites.
Highly accurate protein structure prediction for the human proteome
K. Tunyasuvunakool, J. Adler, Z. Wu, T. Green, M. Zielinski, A. Žídek, A. Bridgland, A. Cowie, C. Meyer, A. Laydon, S. Velankar, G. Kleywegt, A. Bateman, R. Evans, A. Pritzel, M. Figurnov, O. Ronneberger, R. Bates, S. Kohl, and D. Hassabis · 2021
Later among the works it cites.
Tensor Programs IV: Feature learning in infinite-width neural networks
G. Yang and E. J. Hu · 2021
Later among the works it cites.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
E. Abbe, E. Boix-Adsera, and T. Misiakiewicz · 2022
Closest in time.
Neural networks as kernel learners: The silent alignment effect
A. Atanasov, B. Bordelon, and C. Pehlevan · 2022
Closest in time.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
J. Ba, M. A. Erdogdu, T. Suzuki, Z. Wang, D. Wu, and G. Yang · 2022
Closest in time.
Learning single-index models with shallow neural networks
A. Bietti, J. Bruna, C. Sanford, and M. J. Song · 2022
Closest in time.
Self-consistent dynamical field theory of kernel evolution in wide neural networks
B. Bordelon and C. Pehlevan · 2022
Closest in time.
Neural networks can learn representations with gradient descent
A. Damian, J. Lee, and M. Soltanolkotabi · 2022
Closest in time.
Why do tree-based models still outperform deep learning on typical tabular data?
L. Grinsztajn, E. Oyallon, and G. Varoquaux · 2022
Closest in time.
Grokking: Generalization beyond overfitting on small algorithmic datasets
A. Power, Y. Burda, H. Edwards, I. Babuschkin, and V. Misra · 2022
Closest in time.
The Principles of Deep Learning Theory: An Effective Theory Approach to Understanding Neural Networks
D. A. Roberts, S. Yaida, and B. Hanin · 2022
Closest in time.
Revisiting pretraining objectives for tabular deep learning
I. Rubachev, A. Alekberov, Y. Gorishniy, and A. Babenko · 2022
Closest in time.
A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features
Z. Shi, J. Wei, and Y. Lian · 2022
Closest in time.
Salient ImageNet: How to discover spurious features in deep learning?
S. Singla and S. Feizi · 2022
Closest in time.
Limitations of the ntk for understanding generalization in deep learning
N. Vyas, Y. Bansal, and P. Nakkiran · 2022
Closest in time.
Quadratic models for understanding neural network dynamics
L. Zhu, C. Liu, A. Radhakrishnan, and M. Belkin · 2022
Closest in time.
Neural networks efficiently learn low-dimensional representations with sgd
A. Mousavi=Hosseini, S. Park, M. Girotti, I. Mitliagkas, and M. A. Erdogu · 2023
Closest in time.