Fetching the paper…
Reading the bibliography…
Generalization performance of classifiers in deep learning has recently become a subject of intense study.
Theory of reproducing kernels
Nachman Aronszajn · 1950
Earlier work this paper cites.
Nearest neighbor pattern classification
Thomas Cover and Peter Hart · 1967
Earlier work this paper cites.
Darpa timit acoustic-phonetic continous speech corpus cd-rom
John S Garofolo, Lori F Lamel, William M Fisher, Jonathon G Fiscus, and David S Pallett · 1993
Earlier work this paper cites.
Reflections after refereeing papers for nips
Leo Breiman · 1995
Earlier work this paper cites.
Newsweeder: Learning to filter netnews
Ken Lang · 1995
Earlier work this paper cites.
The Nature of Statistical Learning Theory
Vladimir N. Vapnik · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Boosting the margin: a new explanation for the effectiveness of voting methods
Robert E. Schapire, Yoav Freund, Peter Bartlett, and Wee Sun Lee · 1998
Earlier work this paper cites.
Svm vs regularized least squares classification
P Zhang and Jing Peng · 2004
Earlier work this paper cites.
On early stopping in gradient descent learning
Yuan Yao, Lorenzo Rosasco, and Andrea Caponnetto · 2007
Earlier work this paper cites.
Support vector machines
Ingo Steinwart and Andreas Christmann · 2008
Earlier work this paper cites.
Neural network learning: Theoretical foundations
Martin Anthony and Peter L Bartlett · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Ng · 2011
Earlier work this paper cites.
Pegasos: Primal estimated sub-gradient solver for svm
Shai Shalev-Shwartz, Yoram Singer, Nathan Srebro, and Andrew Cotter · 2011
Cited alongside, same era.
An algorithm to improve speech recognition in noise for hearing-impaired listeners
Eric W Healy, Sarah E Yoho, Yuxuan Wang, and DeLiang Wang · 2013
Cited alongside, same era.
Mini-batch primal and dual methods for svms
Martin Takác, Avleen Singh Bijral, Peter Richtárik, and Nati Srebro · 2013
Cited alongside, same era.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus · 2013
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Spectrally-normalized margin bounds for neural networks
Peter Bartlett, Dylan J Foster, and Matus Telgarsky · 2017
Later among the works it cites.
Eigenvalue decay implies polynomial-time learnability for neural networks
Surbhi Goel and Adam R. Klivans · 2017
Later among the works it cites.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso A. Poggio, Alexander Rakhlin, and James Stokes · 2017
Later among the works it cites.
Diving into the shallows: a computational perspective on large-scale shallow learning
Siyuan Ma and Mikhail Belkin · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Cited alongside, same era.
Early stopping and non-parametric regression: an optimal data-dependent stopping rule
Garvesh Raskutti, Martin J Wainwright, and Bin Yu · 2014
Cited alongside, same era.
Less is more: Nyström computational regularization
Alessandro Rudi, Raffaello Camoriano, and Lorenzo Rosasco · 2015
Cited alongside, same era.
NYTRO: When subsampling meets early stopping
R. Camoriano, T. Angles, A. Rudi, and L. Rosasco · 2016
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, and Yann LeCun · 2016
Cited alongside, same era.
An analysis of deep neural network models for practical applications
Alfredo Canziani, Adam Paszke, and Eugenio Culurciello · 2016
Cited alongside, same era.
Solving ridge regression using sketched preconditioned svrg
Alon Gonen, Francesco Orabona, and Shai Shalev-Shwartz · 2016
Cited alongside, same era.
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2017
Later among the works it cites.
Kernel approximation methods for speech recognition
Avner May, Alireza Bagheri Garakani, Zhiyun Lu, Dong Guo, Kuan Liu, Aurélien Bellet, Linxi Fan, Michael Collins, Daniel Hsu, Brian Kingsbury, et al · 2017
Later among the works it cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Later among the works it cites.
Theory of Deep Learning III: explaining the non-overfitting puzzle
T. Poggio, K. Kawaguchi, Q. Liao, B. Miranda, L. Rosasco, X. Boix, J. Hidary, and H. Mhaskar · 2017
Later among the works it cites.
FALKON: An Optimal Large Scale Kernel Method
A. Rudi, L. Carratino, and L. Rosasco · 2017
Later among the works it cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Later among the works it cites.
The Implicit Bias of Gradient Descent on Separable Data
D. Soudry, E. Hoffer, and N. Srebro · 2017
Later among the works it cites.
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2017
Later among the works it cites.
Explaining the success of adaboost and random forests as interpolating classifiers
Abraham J Wyner, Matthew Olson, Justin Bleich, and David Mease · 2017
Later among the works it cites.
Approximation beats concentration? an approximation view on inference with smooth radial kernels
Mikhail Belkin · 2018
Closest in time.