Fetching the paper…
Reading the bibliography…
We investigate the training dynamics of two-layer neural networks when learning multi-index target functions.
Dynamic theory of the spin-glass phase
H. Sompolinsky and A. Zippelius · 1981
Earlier work this paper cites.
Chaos in random neural networks
H. Sompolinsky, A. Crisanti, and H. J. Sommers · 1988
Earlier work this paper cites.
New method for studying the dynamics of disordered spin systems without finite-size effects
H. Eissfeller and M. Opper · 1992
Earlier work this paper cites.
Mean-field Monte Carlo approach to the Sherrington-Kirkpatrick model with asymmetric couplings
H. Eissfeller and M. Opper · 1994
Earlier work this paper cites.
On-line learning in soft committee machines
D. Saad and S. A. Solla · 1995
Earlier work this paper cites.
Dynamical mean-field theory of strongly correlated fermion systems and the limit of infinite dimensions
A. Georges, G. Kotliar, W. Krauth, and M. J. Rozenberg · 1996
Earlier work this paper cites.
Symmetric langevin spin glass dynamics
G. Ben Arous, A. Guionnet, et al · 1997
Earlier work this paper cites.
Out of equilibrium dynamics in spin-glasses and other glassy systems
J.-P. Bouchaud, L. F. Cugliandolo, J. Kurchan, and M. Mézard · 1998
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
M. Kearns · 1998
Earlier work this paper cites.
Statistical mechanics of support vector networks
R. Dietrich, M. Opper, and H. Sompolinsky · 1999
Earlier work this paper cites.
Dynamics of glassy systems
L. F. Cugliandolo · 2003
Earlier work this paper cites.
Special functions
G. E. Andrews · 2004
Earlier work this paper cites.
The dynamics of message passing on dense graphs, with applications to compressed sensing
M. Bayati and A. Montanari · 2011
Earlier work this paper cites.
An iterative construction of solutions of the tap equations for the sherrington–kirkpatrick model
E. Bolthausen · 2014
Earlier work this paper cites.
Out-of-equilibrium dynamical mean-field equations for the perceptron model
E. Agoritsas, G. Biroli, P. Urbani, and F. Zamponi · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P.-M. Nguyen · 2018
Earlier work this paper cites.
The committee machine: computational to statistical gaps in learning a two-layers neural network
B. Aubin, A. Maillard, J. Barbier, F. Krzakala, N. Macris, and L. Zdeborová · 2019
Earlier work this paper cites.
Optimal errors and phase transitions in high-dimensional generalized linear models
J. Barbier, F. Krzakala, N. Macris, L. Miolane, and L. Zdeborová · 2019
Earlier work this paper cites.
Limitations of lazy training of two-layers neural network
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Cited alongside, same era.
Numerical implementation of dynamical mean field theory for disordered systems: application to the lotka–volterra model of ecosystems
F. Roy, G. Biroli, G. Bunin, and C. Cammarota · 2019
Cited alongside, same era.
Spectrum dependent learning curves in kernel regression and wide neural networks
B. Bordelon, A. Canatar, and C. Pehlevan · 2020
Cited alongside, same era.
Learning polynomials in few relevant dimensions
S. Chen and R. Meka · 2020
Cited alongside, same era.
Algorithms and sq lower bounds for pac learning one-hidden-layer relu networks
I. Diakonikolas, D. M. Kane, V. Kontonis, and N. Zarifis · 2020
Cited alongside, same era.
When do neural networks outperform kernel methods?
Stochasticity helps to navigate rough landscapes: comparing gradient-descent-based algorithms in the phase retrieval problem
F. Mignacco, P. Urbani, and L. Zdeborová · 2021
Later among the works it cites.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
E. Abbe, E. Boix-Adsera, and T. Misiakiewicz · 2022
Later among the works it cites.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
J. Ba, M. A. Erdogdu, T. Suzuki, Z. Wang, D. Wu, and G. Yang · 2022
Later among the works it cites.
High-dimensional limit theorems for sgd: Effective dynamics and critical scaling
G. Ben Arous, R. Gheissari, and A. Jagannath · 2022
Later among the works it cites.
Hardness of noise-free learning for two-hidden-layer neural networks
S. Chen, A. Gollakota, A. Klivans, and R. Meka · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2020
Cited alongside, same era.
Superpolynomial lower bounds for learning one-layer neural networks using gradient descent
S. Goel, A. Gollakota, Z. Jin, S. Karmalkar, and A. Klivans · 2020
Cited alongside, same era.
Phase retrieval in high dimensions: Statistical and computational phase transitions, 2020
A. Maillard, B. Loureiro, F. Krzakala, and L. Zdeborová · 2020
Cited alongside, same era.
Marvels and pitfalls of the langevin algorithm in noisy high-dimensional inference
S. S. Mannelli, G. Biroli, C. Cammarota, F. Krzakala, P. Urbani, and L. Zdeborová · 2020
Cited alongside, same era.
Dynamical mean-field theory for stochastic gradient descent in gaussian mixture classification
F. Mignacco, F. Krzakala, P. Urbani, and L. Zdeborová · 2020
Cited alongside, same era.
Mean field analysis of neural networks: A central limit theorem
J. Sirignano and K. Spiliopoulos · 2020
Cited alongside, same era.
The staircase property: How hierarchical structure can guide deep learning
E. Abbe, E. Boix-Adsera, M. S. Brennan, G. Bresler, and D. Nagaraj · 2021
Cited alongside, same era.
Neural networks can learn representations with gradient descent
A. Damian, J. Lee, and M. Soltanolkotabi · 2022
Later among the works it cites.
The effective noise of stochastic gradient descent
F. Mignacco and P. Urbani · 2022
Later among the works it cites.
Universality of empirical risk minimization
A. Montanari and B. N. Saeed · 2022
Later among the works it cites.
Trainability and accuracy of artificial neural networks: An interacting particle system approach
G. Rotskoff and E. Vanden-Eijnden · 2022
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics, 2023
E. Abbe, E. Boix-Adsera, and T. Misiakiewicz · 2023
Later among the works it cites.
Learning in the presence of low-dimensional structure: a spiked random matrix perspective
J. Ba, M. A. Erdogdu, T. Suzuki, Z. Wang, and D. Wu · 2023
Later among the works it cites.
On learning gaussian multi-index models with gradient flow
A. Bietti, J. Bruna, and L. Pillaud-Vivien · 2023
Later among the works it cites.
A. Damian, E. Nichani, R. Ge, and J. D. Lee · 2023
Later among the works it cites.
How two-layer neural networks learn, one (giant) step at a time, 2023
Y. Dandi, F. Krzakala, B. Loureiro, L. Pesce, and L. Stephan · 2023
Later among the works it cites.
Rigorous dynamical mean field theory for stochastic gradient descent methods, 2023
C. Gerbelot, E. Troiani, F. Mignacco, F. Krzakala, and L. Zdeborova · 2023
Later among the works it cites.
A theory of non-linear feature learning with one gradient step in two-layer neural networks, 2023
B. Moniri, D. Lee, H. Hassani, and E. Dobriban · 2023
Later among the works it cites.
Gradient-based feature learning under structured data, 2023
A. Mousavi-Hosseini, D. Wu, T. Suzuki, and M. A. Erdogdu · 2023
Later among the works it cites.
Symmetric single index learning, 2023
A. Zweig and J. Bruna · 2023
Later among the works it cites.