Fetching the paper…
Reading the bibliography…
We study gradient flow on the multi-index regression problem for high-dimensional Gaussian data.
Towards understanding the spectral bias of deep learning
Cao, Y., Fang, Z., Wu, Y., Zhou, D.-X., and Gu, Q. (2019) · 1912
Earlier work this paper cites.
Learning single-index models in gaussian space
Dudeja, R. and Hsu, D. (2018) · 1930
Earlier work this paper cites.
Zufällige bewegungen
Kolmogorov, A. N. (1934) · 1934
Earlier work this paper cites.
Orthogonal polynomials
Szego, G. (1939) · 1939
Earlier work this paper cites.
Some elementary inequalities relating to the gamma and incomplete gamma function
Gautschi, W. (1959) · 1959
Earlier work this paper cites.
Sliced inverse regression for dimension reduction
Li, K.-C. (1991) · 1991
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Barron, A. R. (1993) · 1993
Earlier work this paper cites.
On the ornstein-uhlenbeck operator in spaces of continuous functions
Daprato, G. and Lunardi, A. (1995) · 1995
Earlier work this paper cites.
The geometry of algorithms with orthogonality constraints
Edelman, A., Arias, T. A., and Smith, S. T. (1998) · 1998
Earlier work this paper cites.
Approximation bounds for smooth functions in c (r/sup d/) by neural and mixture networks
Maiorov, V. and Meir, R. S. (1998) · 1998
Earlier work this paper cites.
Structure adaptive approach for dimension reduction
Hristache, M., Juditsky, A., Polzehl, J., and Spokoiny, V. (2001) · 2001
Earlier work this paper cites.
An introduction to Morse theory
Matsumoto, Y. (2002) · 2002
Earlier work this paper cites.
Envelope theorems for arbitrary choice sets
Milgrom, P. and Segal, I. (2002) · 2002
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Schölkopf, B. and Smola, A. (2002) · 2002
Earlier work this paper cites.
An adaptive estimation of dimension reduction space
Xia, Y., Tong, H., Li, W. K., and Zhu, L.-X. (2002) · 2002
Earlier work this paper cites.
Statistics on special manifolds
Chikuse, Y. (2003) · 2003
Earlier work this paper cites.
On the largest principal angle between random subspaces
Absil, P.-A., Edelman, A., and Koev, P. (2006) · 2006
Earlier work this paper cites.
Noise-induced phenomena in slow-fast dynamical systems: a sample-paths approach
Berglund, N. and Gentz, B. (2006) · 2006
Earlier work this paper cites.
Wiener chaos expansion and numerical solutions of stochastic partial differential equations
Luo, W. (2006) · 2006
Earlier work this paper cites.
Kernel techniques: from machine learning to meshless methods
Schaback, R. and Wendland, H. (2006) · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B. (2007) · 2007
Earlier work this paper cites.
The isotron algorithm: High-dimensional isotonic regression
Kalai, A. T. and Sastry, R. (2009) · 2009
Earlier work this paper cites.
Aspects of multivariate statistical theory
Muirhead, R. J. (2009) · 2009
Earlier work this paper cites.
Learning kernel-based halfspaces with the zero-one loss
Shalev-Shwartz, S., Shamir, O., and Sridharan, K. (2010) · 2010
Earlier work this paper cites.
Noisy heteroclinic networks
Bakhtin, Y. (2011) · 2011
Earlier work this paper cites.
A grassmann manifold handbook: Basic geometry and computational aspects
Bendokat, T., Zimmermann, R., and Absil, P.-A. (2020) · 2011
Earlier work this paper cites.
Efficient learning of generalized linear and single index models with isotonic regression
Kakade, S. M., Kanade, V., Shamir, O., and Kalai, A. (2011) · 2011
Cited alongside, same era.
Spectral theory of self-adjoint operators in Hilbert space
Birman, M. S. and Solomjak, M. Z. (2012) · 2012
Cited alongside, same era.
Matrix analysis
Horn, R. A. and Johnson, C. R. (2012) · 2012
Cited alongside, same era.
Li, Z., Luo, Y., and Lyu, K. (2020) · 2012
Cited alongside, same era.
On some inequalities for the gamma function
Laforgia, A. and Natalini, P. (2013) · 2013
Cited alongside, same era.
Angles between subspaces and their tangents
Zhu, P. and Knyazev, A. (2013) · 2013
Prediction under latent factor regression: Adaptive pcr, interpolating predictors and beyond
Bing, X., Bunea, F., Strimas-Mackey, S., and Wegkamp, M. (2021) · 2021
Later among the works it cites.
Jacot, A., Ged, F., Şimşek, B., Hongler, C., and Gabriel, F. (2021) · 2021
Later among the works it cites.
On the cryptographic hardness of learning single periodic neurons
Song, M. J., Zadik, I., and Bruna, J. (2021) · 2021
Later among the works it cites.
Abbe, E., Boix-Adsera, E., and Misiakiewicz, T. (2022) · 2022
Later among the works it cites.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
High-dimensional integration on rd, weighted hermite spaces, and orthogonal transforms
Irrgeher, C. and Leobacher, G. (2015) · 2015
Cited alongside, same era.
Ridge functions
Pinkus, A. (2015) · 2015
Cited alongside, same era.
Gradient descent only converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B. (2016) · 2016
Cited alongside, same era.
The landscape of empirical risk for non-convex losses
Mei, S., Bai, Y., and Montanari, A. (2016) · 2016
Cited alongside, same era.
Differentiating the singular value decomposition
Townsend, J. (2016) · 2016
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Bach, F. (2017) · 2017
Cited alongside, same era.
Ba, J., Erdogdu, M. A., Suzuki, T., Wang, Z., Wu, D., and Yang, G. (2022) · 2022
Later among the works it cites.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Barak, B., Edelman, B., Goel, S., Kakade, S., Malach, E., and Zhang, C. (2022) · 2022
Later among the works it cites.
High-dimensional limit theorems for sgd: Effective dynamics and critical scaling
Ben Arous, G., Gheissari, R., and Jagannath, A. (2022) · 2022
Later among the works it cites.
Learning single-index models with shallow neural networks
Bietti, A., Bruna, J., Sanford, C., and Song, M. J. (2022) · 2022
Later among the works it cites.
Neural networks can learn representations with gradient descent
Damian, A., Lee, J., and Soltanolkotabi, M. (2022) · 2022
Later among the works it cites.
Learning a single neuron for non-monotonic activation functions
Wu, L. (2022) · 2022
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Abbe, E., Boix-Adsera, E., and Misiakiewicz, T. (2023) · 2023
Closest in time.
Arnaboldi, L., Stephan, L., Krzakala, F., and Loureiro, B. (2023) · 2023
Closest in time.
Incremental learning in diagonal linear networks
Berthier, R. (2023) · 2023
Closest in time.
Learning time-scales in two-layers neural networks
Berthier, R., Montanari, A., and Zhou, K. (2023) · 2023
Closest in time.
On single index models beyond gaussian data
Bruna, J., Pillaud-Vivien, L., and Zweig, A. (2023) · 2023
Closest in time.
Learning narrow one-hidden-layer relu networks
Chen, S., Dou, Z., Goel, S., Klivans, A. R., and Meka, R. (2023) · 2023
Closest in time.
Damian, A., Nichani, E., Ge, R., and Lee, J. D. (2023) · 2023
Closest in time.
Learning two-layer neural networks, one (giant) step at a time
Dandi, Y., Krzakala, F., Loureiro, B., Pesce, L., and Stephan, L. (2023) · 2023
Closest in time.
Nonparametric linear feature learning in regression through regularisation
Follain, B., Simsekli, U., and Bach, F. (2023) · 2023
Closest in time.
Glasgow, M. (2023) · 2023
Closest in time.
Mahankali, A., Haochen, J. Z., Dong, K., Glasgow, M., and Ma, T. (2023) · 2023
Closest in time.
Leveraging the two timescale regime to demonstrate convergence of neural networks
Marion, P. and Berthier, R. (2023) · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Liberum, T., Smith, J., and Steinhardt, J. (2023) · 2023
Closest in time.
Saddle-to-saddle dynamics in diagonal linear networks
Pesme, S. and Flammarion, N. (2023) · 2023
Closest in time.
Symmetric single index learning
Zweig, A. and Bruna, J. (2023) · 2023
Closest in time.