Fetching the paper…
Reading the bibliography…
Single-index models are a class of functions given by an unknown univariate ``link'' function applied to an unknown one-dimensional projection of the input.
The free markoff field
Edward Nelson · 1973
Earlier work this paper cites.
A projection pursuit algorithm for exploratory data analysis
Jerome H Friedman and John W Tukey · 1974
Earlier work this paper cites.
Logarithmic sobolev inequalities
Leonard Gross · 1975
Earlier work this paper cites.
Projection pursuit
Peter J Huber · 1985
Earlier work this paper cites.
Sliced inverse regression for dimension reduction
Ker-Chau Li · 1991
Earlier work this paper cites.
On principal hessian directions for data visualization and dimension reduction: Another application of stein’s lemma
Ker-Chau Li · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Semiparametric least squares (sls) and weighted sls estimation of single-index models
Hidehiko Ichimura · 1993
Earlier work this paper cites.
Geometric categories and o-minimal structures
Lou Van den Dries and Chris Miller · 1996
Earlier work this paper cites.
Integral transforms, reproducing kernels and their applications
Saburou Saitoh · 1997
Earlier work this paper cites.
Approximation theory of the mlp model in neural networks
Allan Pinkus · 1999
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
Direct estimation of the index coefficient in a single-index model
Marian Hristache, Anatoli Juditsky, and Vladimir Spokoiny · 2001
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Bernhard Schölkopf and Alexander Smola · 2002
Earlier work this paper cites.
Nonparametric and semiparametric models
Wolfgang Härdle, Marlene Müller, Stefan Sperlich, and Axel Werwatz · 2004
Earlier work this paper cites.
Gradient flows: in metric spaces and in the space of probability measures
Luigi Ambrosio, Nicola Gigli, and Giuseppe Savaré · 2005
Earlier work this paper cites.
Convex Analysis
Jonathan Borwein and Adrian Lewis · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
A new algorithm for estimating the effective dimension-reduction subspace
Arnak S Dalalyan, Anatoly Juditsky, and Vladimir Spokoiny · 2008
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence Saul · 2009
Earlier work this paper cites.
Reproducing kernel Hilbert spaces in probability and statistics
Alain Berlinet and Christine Thomas-Agnan · 2011
Earlier work this paper cites.
Chernoff-type bounds for the gaussian error function
Seok-Ho Chang, Pamela C Cosman, and Laurence B Milstein · 2011
Earlier work this paper cites.
User-friendly tail bounds for sums of random matrices
Joel A Tropp · 2012
Earlier work this paper cites.
Analysis of boolean functions
Ryan O’Donnell · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Curves of descent
Dmitriy Drusvyatskiy, Alexander D Ioffe, and Adrian S Lewis · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
The non-convex burer-monteiro approach works on smooth semidefinite programs
Nicolas Boumal, Vlad Voroninski, and Afonso Bandeira · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
The landscape of empirical risk for non-convex losses
Song Mei, Yu Bai, and Andrea Montanari · 2016
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Cited alongside, same era.
On the equivalence between kernel quadrature rules and random feature expansions
Francis Bach · 2017
Cited alongside, same era.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2017
Cited alongside, same era.
On the optimization landscape of tensor decompositions
Rong Ge and Tengyu Ma · 2017
Cited alongside, same era.
How to escape saddle points efficiently, 2017
Towards understanding hierarchical learning: Benefits of neural representations
Minshuo Chen, Yu Bai, Jason D Lee, Tuo Zhao, Huan Wang, Caiming Xiong, and Richard Socher · 2020
Later among the works it cites.
Stochastic subgradient method converges on tame functions
Damek Davis, Dmitriy Drusvyatskiy, Sham Kakade, and Jason D Lee · 2020
Later among the works it cites.
Algorithms and sq lower bounds for pac learning one-hidden-layer relu networks
Ilias Diakonikolas, Daniel M. Kane, Vasilis Kontonis, and Nikos Zarifis · 2020
Later among the works it cites.
Hardness of learning neural networks with natural weights
Amit Daniely and Gal Vardi · 2020
Later among the works it cites.
Superpolynomial lower bounds for learning one-layer neural networks using gradient descent
Surbhi Goel, Aravind Gollakota, Zhihan Jin, Sushrut Karmalkar, and Adam Klivans · 2020
Later among the works it cites.
Directional convergence and alignment in deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Cited alongside, same era.
On some extensions of bernstein’s inequality for self-adjoint operators
Stanislav Minsker · 2017
Cited alongside, same era.
Probability and computing: Randomization and probabilistic techniques in algorithms and data analysis
Michael Mitzenmacher and Eli Upfal · 2017
Cited alongside, same era.
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco · 2017
Cited alongside, same era.
Learning relus via gradient descent
Mahdi Soltanolkotabi · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Cited alongside, same era.
Double/debiased machine learning for treatment and structural parameters
Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins · 2018
Cited alongside, same era.
Ziwei Ji and Matus Telgarsky · 2020
Later among the works it cites.
Marvels and pitfalls of the langevin algorithm in noisy high-dimensional inference
Stefano Sarao Mannelli, Giulio Biroli, Chiara Cammarota, Florent Krzakala, Pierfrancesco Urbani, and Lenka Zdeborová · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
Learning theory from first principles, 2021
Francis Bach · 2021
Later among the works it cites.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2021
Later among the works it cites.
Numerical influence of relu’(0) on backpropagation
David Bertoin, Jérôme Bolte, Sébastien Gerchinovitz, and Edouard Pauwels · 2021
Later among the works it cites.
Prediction under latent factor regression: Adaptive pcr, interpolating predictors and beyond
Xin Bing, Florentina Bunea, Seth Strimas-Mackey, and Marten Wegkamp · 2021
Later among the works it cites.
Oracle complexity in nonsmooth nonconvex optimization
Guy Kornowski and Ohad Shamir · 2021
Later among the works it cites.
Quantifying the benefit of using differentiable learning over tangent kernels
Eran Malach, Pritish Kamath, Emmanuel Abbe, and Nathan Srebro · 2021
Later among the works it cites.
The connection between approximation, depth separation and learnability in neural networks
Eran Malach, Gilad Yehudai, Shai Shalev-Schwartz, and Ohad Shamir · 2021
Later among the works it cites.
Optimization-based separations for neural networks
Itay Safran and Jason D Lee · 2021
Later among the works it cites.
On the cryptographic hardness of learning single periodic neurons
Min Jae Song, Ilias Zadik, and Joan Bruna · 2021
Later among the works it cites.
A local convergence theory for mildly over-parameterized two-layer neural network
Mo Zhou, Rong Ge, and Chi Jin · 2021
Later among the works it cites.
Emmanuel Abbe, Enric Boix-Adsera, and Theodor Misiakiewicz · 2022
Closest in time.
High-dimensional limit theorems for sgd: Effective dynamics and critical scaling
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2022
Closest in time.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Closest in time.
Hardness of noise-free learning for two-hidden-layer neural networks
Sitan Chen, Aravind Gollakota, Adam R Klivans, and Raghu Meka · 2022
Closest in time.
http://dlmf.nist.gov/, Release 1.1.6 of 2022-06-30
NIST Digital Library of Mathematical Functions · 2022
Closest in time.
Neural networks can learn representations with gradient descent
Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi · 2022
Closest in time.
Neural networks efficiently learn low-dimensional representations with sgd
Alireza Mousavi-Hosseini, Sejun Park, Manuela Girotti, Ioannis Mitliagkas, and Murat A Erdogdu · 2022
Closest in time.
Generalization error of random feature and kernel methods: Hypercontractivity and kernel matrix concentration
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2022
Closest in time.
Eshaan Nichani, Yu Bai, and Jason D Lee · 2022
Closest in time.
Phase diagram of stochastic gradient descent in high-dimensional two-layer neural networks
Rodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2022
Closest in time.