Fetching the paper…
Reading the bibliography…
Neural networks extract features from data using stochastic gradient descent (SGD).
On-line backpropagation in two-layered neural networks
Riegler, P. and Biehl, M · 1995
Earlier work this paper cites.
Statistical mechanics of support vector networks
Dietrich, R., Opper, M., and Sompolinsky, H · 1999
Earlier work this paper cites.
Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices
Baik, J., Arous, G. B., and Péché, S · 2005
Earlier work this paper cites.
Tensor decompositions for learning latent variable models
Anandkumar, A., Ge, R., Hsu, D., Kakade, S. M., and Telgarsky, M · 2014
Earlier work this paper cites.
A statistical model for tensor pca
Richard, E. and Montanari, A · 2014
Earlier work this paper cites.
Bayesian estimation from few samples: com- munity detection and related problems
Hopkins, S. B. and Steurer, D · 2017
Earlier work this paper cites.
The power of sum-of-squares for detecting hidden structures
Hopkins, S. B., Kothari, P. K., Potechin, A., Raghavendra, P., Schramm, T., and Steurer, D · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Li, Y. and Yuan, Y · 2017
Earlier work this paper cites.
Statistical inference and the sum of squares method
Hopkins, S · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured Data
Li, Y. and Liang, Y · 2018
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., and Wang, R · 2019
Earlier work this paper cites.
A nearly tight sum-of-squares lower bound for the planted clique problem
Barak, B., Hopkins, S., Kelner, J., Kothari, P. K., Moitra, A., and Potechin, A · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Earlier work this paper cites.
Limitations of lazy training of two-layers neural network
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A · 2019
Earlier work this paper cites.
Sgd on neural networks learns functions of increasing complexity
Kalimeris, D., Kaplun, G., Nakkiran, P., Edelman, B., Yang, T., Barak, B., and Zhang, H · 2019
Earlier work this paper cites.
Notes on computational hardness of hypothesis testing: Predictions using the low-degree likelihood ratio
Kunisky, D., Wein, A. S., and Bandeira, A. S · 2019
Earlier work this paper cites.
Spectrum dependent learning curves in kernel regression and wide neural networks
Bordelon, B., Canatar, A., and Pehlevan, C · 2020
Earlier work this paper cites.
Generalisation error in learning with random features and the hidden manifold model
Gerace, F., Loureiro, B., Krzakala, F., Mézard, M., and Zdeborová, L · 2020
Cited alongside, same era.
When do neural networks outperform kernel methods?
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A · 2020
Cited alongside, same era.
Modeling the influence of data structure on learning in neural networks: The hidden manifold model
Goldt, S., Mézard, M., Krzakala, F., and Zdeborová, L · 2020
Cited alongside, same era.
Marvels and pitfalls of the langevin algorithm in noisy high-dimensional inference
Sarao Mannelli, S., Biroli, G., Cammarota, C., Krzakala, F., Urbani, P., and Zdeborová, L · 2020
Cited alongside, same era.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Spigler, S., Geiger, M., and Wyart, M · 2020
Cited alongside, same era.
Data-driven emergence of convolutional structure in neural networks
Ingrosso, A. and Goldt, S · 2022
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Mei, S. and Montanari, A · 2022
Later among the works it cites.
Precise learning curves and higher-order scalings for dot-product kernel regression
Xiao, L., Hu, H., Misiakiewicz, T., Lu, Y., and Pennington, J · 2022
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Abbé, E., Adserà, E. B., and Misiakiewicz, T · 2023
Later among the works it cites.
Learning time-scales in two-layers neural networks
Berthier, R., Montanari, A., and Zhou, K · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abbé, E., Boix-Adsera, E., Brennan, M. S., Bresler, G., and Nagaraj, D · 2021
Cited alongside, same era.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Ben Arous, G., Gheissari, R., and Jagannath, A · 2021
Cited alongside, same era.
Deep equals shallow for reLU networks in kernel regimes
Bietti, A. and Bach, F · 2021
Cited alongside, same era.
Jacot, A., Ged, F., Şimşek, B., Hongler, C., and Gabriel, F · 2021
Cited alongside, same era.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Refinetti, M., Goldt, S., Krzakala, F., and Zdeborová, L · 2021
Cited alongside, same era.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
Abbé, E., Adsera, E. B., and Misiakiewicz, T · 2022
Cited alongside, same era.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Ba, J., Erdogdu, M. A., Suzuki, T., Wang, Z., Wu, D., and Yang, G · 2022
Cited alongside, same era.
Bietti, A., Bruna, J., and Pillaud-Vivien, L · 2023
Later among the works it cites.
Optimal learning of deep random networks of extensive-width
Cui, H., Krzakala, F., and Zdeborová, L · 2023
Later among the works it cites.
Damian, A., Nichani, E., Ge, R., and Lee, J. D · 2023
Later among the works it cites.
Learning two-layer neural networks, one (giant) step at a time
Dandi, Y., Krzakala, F., Loureiro, B., Pesce, L., and Stephan, L · 2023
Later among the works it cites.
Learning interacting theories from data
Merger, C., René, A., Fischer, K., Bouss, P., Nestler, S., Dahmen, D., Honerkamp, C., and Helias, M · 2023
Later among the works it cites.
Statistical temporal pattern extraction by neuronal architecture
Nestler, S., Helias, M., and Gilson, M · 2023
Later among the works it cites.
Pesce, L., Krzakala, F., Loureiro, B., and Stephan, L · 2023
Later among the works it cites.
Neural networks trained with sgd learn distributions of increasing complexity
Refinetti, M., Ingrosso, A., and Goldt, S · 2023
Later among the works it cites.
Learning from higher-order statistics, efficiently: hypothesis tests, random features, and neural networks
Székely, E., Bardone, L., Gerace, F., and Goldt, S · 2023
Later among the works it cites.
Neural networks learn statistics of increasing complexity
Belrose, N., Pope, Q., Quirke, L., Mallen, A., and Fern, X · 2024
Closest in time.
The computational complexity of learning gaussian single-index models
Damian, A., Pillaud-Vivien, L., Lee, J. D., and Bruna, J · 2024
Closest in time.
Dandi, Y., Troiani, E., Arnaboldi, L., Pesce, L., Zdeborová, L., and Krzakala, F · 2024
Closest in time.
Gradient-based feature learning under structured data
Mousavi-Hosseini, A., Wu, D., Suzuki, T., and Erdogdu, M. A · 2024
Closest in time.