Fetching the paper…
Reading the bibliography…
We focus on the task of learning a single index model $\sigma(w^\star \cdot x)$ with respect to the isotropic Gaussian distribution in $d$ dimensions.
Optimum Bounds for the Distributions of Martingales in Banach Spaces
Iosif Pinelis · 1994
Earlier work this paper cites.
Characterizing statistical query learning: Simplified notions and proofs
Balázs Szörényi · 2009
Earlier work this paper cites.
Efficient learning of generalized linear and single index models with isotonic regression
Sham M Kakade, Varun Kanade, Ohad Shamir, and Adam Kalai · 2011
Earlier work this paper cites.
A statistical model for tensor pca
Emile Richard and Andrea Montanari · 2014
Earlier work this paper cites.
Phase retrieval via wirtinger flow: Theory and algorithms
Emmanuel J. Candès, Xiaodong Li, and Mahdi Soltanolkotabi · 2015
Earlier work this paper cites.
Tensor principal component analysis via sum-of-square proofs
Samuel B. Hopkins, Jonathan Shi, and David Steurer · 2015
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Fast spectral algorithms from sum-of-squares proofs: Tensor decomposition and planted sparse vectors
Samuel B. Hopkins, Tselil Schramm, Jonathan Shi, and David Steurer · 2016
Earlier work this paper cites.
Learning relus via gradient descent
Mahdi Soltanolkotabi · 2017
Earlier work this paper cites.
Homotopy analysis for tensor pca
Anima Anandkumar, Yuan Deng, Rong Ge, and Hossein Mobahi · 2017
Earlier work this paper cites.
The landscape of empirical risk for nonconvex losses
Song Mei, Yu Bai, and Andrea Montanari · 2018
Earlier work this paper cites.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2018
Earlier work this paper cites.
A geometric analysis of phase retrieval
Ju Sun, Qing Qu, and John Wright · 2018
Earlier work this paper cites.
Learning single-index models in gaussian space
Rishabh Dudeja and Daniel Hsu · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
Yeming Wen, Kevin Luk, Maxime Gazeau, Guodong Zhang, Harris Chan, and Jimmy Ba · 2019
Cited alongside, same era.
Gradient descent with random initialization: fast global convergence for nonconvex phase retrieval
Yuxin Chen, Yuejie Chi, Jianqing Fan, and Cong Ma · 2019
Cited alongside, same era.
Notes on computational hardness of hypothesis testing: Predictions using the low-degree likelihood ratio, 2019
Dmitriy Kunisky, Alexander S. Wein, and Afonso S. Bandeira · 2019
Cited alongside, same era.
Experiment tracking with weights and biases, 2020
Lukas Biewald · 2020
Later among the works it cites.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2021
Later among the works it cites.
Label noise SGD provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D. Lee · 2021
Later among the works it cites.
Statistical query algorithms and low degree tests are almost equivalent
Matthew S Brennan, Guy Bresler, Sam Hopkins, Jerry Li, and Tselil Schramm · 2021
Later among the works it cites.
Statistical query lower bounds for tensor pca
Rishabh Dudeja and Daniel Hsu · 2021
Later among the works it cites.
Learning single-index models with shallow neural networks
Alberto Bietti, Joan Bruna, Clayton Sanford, and Min Jae Song · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Complex energy landscapes in spiked-tensor and simple glassy models: Ruggedness, arrangements of local minima, and phase transitions
Valentina Ros, Gerard Ben Arous, Giulio Biroli, and Chiara Cammarota · 2019
Cited alongside, same era.
The landscape of the spiked tensor model
Gérard Ben Arous, Song Mei, Andrea Montanari, and Mihai Nica · 2019
Cited alongside, same era.
Why do local methods solve nonconvex problems?, 2020
Tengyu Ma · 2020
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Cited alongside, same era.
Learning polynomials in few relevant dimensions
Sitan Chen and Raghu Meka · 2020
Cited alongside, same era.
How to iron out rough landscapes and get optimal performances: averaged gradient descent and its application to tensor pca
Giulio Biroli, Chiara Cammarota, and Federico Ricci-Tersenghi · 2020
Cited alongside, same era.
Superpolynomial lower bounds for learning one-layer neural networks using gradient descent
Surbhi Goel, Aravind Gollakota, Zhihan Jin, Sushrut Karmalkar, and Adam Klivans · 2020
Cited alongside, same era.
Neural networks can learn representations with gradient descent
Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi · 2022
Later among the works it cites.
What happens after SGD reaches zero loss? –a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2022
Later among the works it cites.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Later among the works it cites.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz · 2022
Later among the works it cites.
The franz-parisi criterion and computational trade-offs in high dimensional statistics
Afonso S Bandeira, Ahmed El Alaoui, Samuel Hopkins, Tselil Schramm, Alexander S Wein, and Ilias Zadik · 2022
Later among the works it cites.
Statistical-computational trade-offs in tensor pca and related problems via communication complexity, 2022
Rishabh Dudeja and Daniel Hsu · 2022
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Emmanuel Abbe, Enric Boix-Adserà, and Theodor Misiakiewicz · 2023
Closest in time.