Fetching the paper…
Reading the bibliography…
We consider the problem of learning function classes computed by neural networks with various activations (e.g.
A theory of the learnable
L. G. Valiant · 1984
Earlier work this paper cites.
Relating data compression and learnability
Nick Littlestone and Manfred Warmuth · 1986
Earlier work this paper cites.
Relating data compression and learnability
Nick Littlestone and Manfred Warmuth · 1986
Earlier work this paper cites.
Decision theoretic generalizations of the pac model for neural net and other learning applications
David Haussler · 1992
Earlier work this paper cites.
Toward efficient agnostic learning
Michael J. Kearns, Robert E. Schapire, and Linda M. Sellie · 1994
Earlier work this paper cites.
Exponentially many local minima for single neurons
Peter Auer, Mark Herbster, and Manfred K. Warmuth · 1996
Earlier work this paper cites.
Generalization bounds via eigenvalues of the gram matrix
B. Schölkopf, J. Shawe-Taylor, AJ. Smola, and RC. Williamson · 1999
Earlier work this paper cites.
Using the nyström method to speed up kernel machines
Christopher KI Williams and Matthias Seeger · 2000
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Bernhard Schölkopf and Alexander J Smola · 2002
Earlier work this paper cites.
Effective dimension and generalization of kernel learning
Tong Zhang · 2003
Earlier work this paper cites.
Local rademacher complexities
Peter L. Bartlett, Olivier Bousquet, and Shahar Mendelson · 2005
Earlier work this paper cites.
On the nyström method for approximating a gram matrix for improved kernel-based learning
Petros Drineas and Michael W Mahoney · 2005
Earlier work this paper cites.
On the eigenspectrum of the gram matrix and the generalization error of kernel-pca
John Shawe-Taylor, Christopher KI Williams, Nello Cristianini, and Jaz Kandola · 2005
Earlier work this paper cites.
Unlabeled compression schemes for maximum classes
Dima Kuzmin and Manfred K. Warmuth · 2007
Earlier work this paper cites.
Relative-error cur matrix decompositions
Petros Drineas, Michael W Mahoney, and S Muthukrishnan · 2008
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Sham M Kakade, Karthik Sridharan, and Ambuj Tewari · 2009
Cited alongside, same era.
Cryptographic hardness for learning intersections of halfspaces
Adam R. Klivans and Alexander A. Sherstov · 2009
Cited alongside, same era.
Finding Cliques in Scale-Free Networks
Anton Krohmer · 2012
Cited alongside, same era.
Learning optimally sparse support vector machines
Andrew Cotter, Shai Shalev-Shwartz, and Nati Srebro · 2013
Cited alongside, same era.
Moment-matching polynomials
Adam R. Klivans and Raghu Meka · 2013
Cited alongside, same era.
Subspace embeddings for the polynomial kernel
Haim Avron, Huy Nguyen, and David Woodruff · 2014
Cited alongside, same era.
Complexity theoretic limitations on learning halfspaces
Amit Daniely · 2016
Later among the works it cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Later among the works it cites.
On statistical learning via the lens of compression
Ofir David, Shay Moran, and Amir Yehudayoff · 2016
Later among the works it cites.
Reliably learning the relu in polynomial time
Surbhi Goel, Varun Kanade, Adam Klivans, and Justin Thaler · 2016
Later among the works it cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Embedding hard learning problems into gaussian space
Adam R. Klivans and Pravesh Kothari · 2014
Cited alongside, same era.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Cited alongside, same era.
Provable methods for training neural networks with sparse connectivity
Hanie Sedghi and Anima Anandkumar · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Matrix coherence and the nystrom method
Ameet Talwalkar and Afshin Rostamizadeh · 2014
Cited alongside, same era.
Algorithmic complexity of power law networks
Pawel Brach, Marek Cygan, Jakub Lacki, and Piotr Sankowski · 2015
Cited alongside, same era.
Cameron Musco and Christopher Musco · 2016
Later among the works it cites.
Distribution-specific hardness of learning neural networks
Ohad Shamir · 2016
Later among the works it cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Later among the works it cites.
Diversity leads to generalization in neural networks
Bo Xie, Yingyu Liang, and Le Song · 2016
Later among the works it cites.
l1-regularized neural networks are improperly learnable in polynomial time
Yuchen Zhang, Jason D Lee, and Michael I Jordan · 2016
Later among the works it cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Closest in time.
SGD learns the conjugate kernel class of the network
Amit Daniely · 2017
Closest in time.
Diving into the shallows: a computational perspective on large-scale shallow learning
Siyuan Ma and Mikhail Belkin · 2017
Closest in time.
On the complexity of learning neural networks
Le Song, Santosh Vempala, John Wilmes, and Bo Xie · 2017
Closest in time.
Electron-proton dynamics in deep learning
Qiuyi Zhang, Rina Panigrahy, and Sushant Sachdeva · 2017
Closest in time.