Fetching the paper…
Reading the bibliography…
In recent years, there are many attempts to understand popular heuristics.
Samet Oymak and Mahdi Soltanolkotabi · 1902
Earlier work this paper cites.
Induction of decision trees
J. Ross Quinlan · 1986
Earlier work this paper cites.
Learning decision lists
Ronald L Rivest · 1987
Earlier work this paper cites.
Learning decision trees from random examples
Andrzej Ehrenfeucht and David Haussler · 1989
Earlier work this paper cites.
Rank-r decision trees are a subclass of r-decision lists
Avrim Blum · 1992
Earlier work this paper cites.
Learning decision trees using the fourier spectrum
Eyal Kushilevitz and Yishay Mansour · 1993
Earlier work this paper cites.
Weakly learning dnf and characterizing statistical query learning using fourier analysis
Avrim Blum, Merrick Furst, Jeffrey Jackson, Michael Kearns, Yishay Mansour, and Steven Rudich · 1994
Earlier work this paper cites.
Boosting theory towards practice: Recent developments in decision tree induction and the weak learning framework
Michael Kearns · 1996
Earlier work this paper cites.
On the boosting ability of top–down decision tree learning algorithms
Michael Kearns and Yishay Mansour · 1999
Earlier work this paper cites.
On using extended statistical queries to avoid membership queries
Nader H Bshouty and Vitaly Feldman · 2002
Earlier work this paper cites.
Noise-tolerant learning, the parity problem, and the statistical query model
Avrim Blum, Adam Kalai, and Hal Wasserman · 2003
Earlier work this paper cites.
On the proper learning of axis-parallel concepts
Nader H Bshouty and Lynn Burroughs · 2003
Earlier work this paper cites.
Learning random log-depth decision trees under the uniform distribution
Jeffrey C Jackson and Rocco A Servedio · 2003
Cited alongside, same era.
Decision trees: More theoretical justification for practical algorithms
Amos Fiat and Dmitry Pechyony · 2004
Cited alongside, same era.
Smoothed analysis of algorithms: Why the simplex algorithm usually takes polynomial time
Daniel A Spielman and Shang-Hua Teng · 2004
Cited alongside, same era.
Learning dnf from random walks
Nader H Bshouty, Elchanan Mossel, Ryan O’Donnell, and Rocco A Servedio · 2005
Cited alongside, same era.
New results for learning noisy parities and halfspaces
Vitaly Feldman, Parikshit Gopalan, Subhash Khot, and Ashok Kumar Ponnuswami · 2006
Cited alongside, same era.
Learning monotone decision trees in polynomial time
Ryan O’Donnell and Rocco A Servedio · 2007
Cited alongside, same era.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Later among the works it cites.
Hyperparameter optimization: A spectral approach
Elad Hazan, Adam Klivans, and Yang Yuan · 2017
Later among the works it cites.
Failures of gradient-based deep learning
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Later among the works it cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Decision trees are pac-learnable from most product distributions: a smoothed analysis
Adam Tauman Kalai and Shang-Hua Teng · 2008
Cited alongside, same era.
On agnostic learning of parities, monomials, and halfspaces
Vitaly Feldman, Parikshit Gopalan, Subhash Khot, and Ashok Kumar Ponnuswami · 2009
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2016
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Cited alongside, same era.
Sitan Chen and Ankur Moitra · 2018
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2018
Later among the works it cites.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.
Chao Ma, Lei Wu, et al · 2019
Closest in time.