Fetching the paper…
Reading the bibliography…
McKernel introduces a framework to use kernel approximates in the mini-batch setting with Stochastic Gradient Descent (SGD) as an alternative to Deep Learning.
A Note on the Generation of Random Normal Deviates
G. E. P. Box and M. E. Muller. 1958 · 1958
Earlier work this paper cites.
The unreasonable effectiveness of mathematics in the natural sciences
E. Wigner. 1960 · 1960
Earlier work this paper cites.
Support vector networks
C. Cortes and V. Vapnik. 1995 · 1995
Earlier work this paper cites.
Regularization theory and neural networks architectures
F. Girosi, M. Jones, and T. Poggio. 1995 · 1995
Earlier work this paper cites.
An equivalence between sparse approximation and Support Vector Machines
F. Girosi. 1998 · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner. 1998 · 1998
Earlier work this paper cites.
In search of the optimal Walsh-Hadamard Transform
J. Johnson and M. Püschel. 2000 · 2000
Earlier work this paper cites.
On the mathematical foundations of learning
F. Cucker and S. Smale. 2001 · 2001
Earlier work this paper cites.
The Mathematics of Learning: Dealing with Data
T. Poggio and S. Smale. 2003 · 2003
Earlier work this paper cites.
The Gray-Code Filter Kernels
G. Ben-Artzi, H. Hel-Or, and Y. Hel-Or. 2007 · 2007
Earlier work this paper cites.
Random Features for Large-Scale Kernel Machines
A. Rahimi and B. Recht. 2007 · 2007
Earlier work this paper cites.
Kernel Methods for Deep Learning
Y. Cho and L. K. Saul. 2009 · 2009
Earlier work this paper cites.
A new learning paradigm: Learning using privileged information
V. Vapnik and A. Vashist. 2009 · 2009
Earlier work this paper cites.
Fast Algorithm for Walsh Hadamard Transform on Sliding Windows
W. Ouyang and W. K. Cham. 2010 · 2010
Earlier work this paper cites.
Fastfood - Approximating Kernel Expansions in Loglinear Time
Q. Le, T. Sarlós, and A. Smola. 2013 · 2013
Earlier work this paper cites.
Faster Ridge Regression via the Subsampled Randomized Hadamard Transform
Y. Lu, P. S. Dhillon, D. Foster, and L. Ungar. 2013 · 2013
Earlier work this paper cites.
Rectifier Nonlinearities Improve Neural Network Acoustic Models
A. L. Maas, A. Y. Hannun, and A. Y. Ng. 2013 · 2013
Earlier work this paper cites.
Learning to Rank Using Privileged Information
V. Sharmanska, N. Quadrianto, and C. H. Lampert. 2013 · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. 2014 · 2014
Cited alongside, same era.
À la Carte - Learning Fast Kernels
Z. Yang, A. Smola, L. Song, and A. G. Wilson. 2014 · 2014
Cited alongside, same era.
Practical and Optimal LSH for Angular Distance
A. Andoni, P. Indyk, T. Laarhoven, I. Razenshteyn, and L. Schmidt. 2015 · 2015
Cited alongside, same era.
Neural Machine Translation by Jointly Learning to Align and Translate
D. Bahdanau, K. Cho, and Y. Bengio. 2015 · 2015
Cited alongside, same era.
Fast Two-Sample Testing with Analytic Representations of Probability Measures
K. Chwialkowski, A. Ramdas, D. Sejdinovic, and A. Gretton. 2015 · 2015
Cited alongside, same era.
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
K. He, X. Zhang, S. Ren, and J. Sun. 2015 · 2015
Self-Normalizing Neural Networks
G. Klarbauer, T. Unterthiner, and A. Mayr. 2017 · 2017
Closest in time.
Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles
B. Lakshminarayanan, A. Pritzel, and C. Blundell. 2017 · 2017
Closest in time.
Theory II: Landscape of the Empirical Risk in Deep Learning
Q. Liao and T. Poggio. 2017 · 2017
Closest in time.
When and Why Are Deep Networks Better than Shallow Ones?
H. Mhaskar, Q. Liao, and T. Poggio. 2017 · 2017
Closest in time.
Why and when can deep-but not shallow-networks avoid the curse of dimensionality: A review
T. Poggio, H. Mhaskar, L. Rosasco, B. Miranda, and Q. Liao. 2017 · 2017
Closest in time.
Generalization Properties of Learning with Random Features
A. Rudi and L. Rosasco. 2017 · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy. 2015 · 2015
Cited alongside, same era.
Effective Approaches to Attention-based Neural Machine Translation
M. Luong, H. Pham, and C. D. Manning. 2015 · 2015
Cited alongside, same era.
On the high-dimensional power of a linear-time two sample test under mean-shift alternatives
S. Reddi, A. Ramdas, A. Singh, B. Poczos, and L. Wasserman. 2015 · 2015
Cited alongside, same era.
Learning Using Privileged Information: Similarity Control and Knowledge Transfer
V. Vapnik and R. Izmailov. 2015 · 2015
Cited alongside, same era.
Classifier Learning with Hidden Information
Z. Wang and Q. Ji. 2015 · 2015
Cited alongside, same era.
Deep Fried Convnets
Z. Yang, M. Moczulski, M. Denil, N. Freitas, A. Smola, L. Song, and Z. Wang. 2015 · 2015
Cited alongside, same era.
A Bayesian Data Augmentation Approach for Learning Deep Models
T. Tran, T. Pham, G. Carneiro, L. Palmer, and I. Reid. 2017 · 2017
Closest in time.
Attention Is All You Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. 2017 · 2017
Closest in time.
Large Margin Object Tracking With Circulant Feature Maps
M. Wang, Y. Liu, and Z. Huang. 2017 · 2017
Closest in time.
FASHION-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
H. Xiao, K. Rasul, and R. Vollgraf. 2017 · 2017
Closest in time.
AutoAugment: Learning Augmentation Policies from Data
E. D. Cubuk, B. Zoph, D. Mane, V. Vasudevan, and Q. V. Le. 2018 · 2018
Closest in time.
Deep Semi-Random Features for Nonlinear Function Approximation
K. Kawaguchi, B. Xie, V. Verma, and L. Song. 2018 · 2018
Closest in time.
Adding One Neuron Can Eliminate All Bad Local Minima
S. Liang, R. Sun, J. D. Lee, and R. Srikant. 2018 · 2018
Closest in time.
Rethinking statistical learning theory: learning using statistical invariants
V. Vapnik and R. Izmailov. 2018 · 2018
Closest in time.
Deep Neural Networks with Multi-Branch Architectures Are Less Non-Convex
H. Zhang, J. Shao, and R. Salakhutdinov. 2018 · 2018
Closest in time.
Elimination of All Bad Local Minima in Deep Learning
K. Kawaguchi and L. P. Kaelbling. 2019 · 2019
Closest in time.
Eliminating all bad Local Minima from Loss Landscapes without even adding an Extra Unit
J. Sohl-Dickstein and K. Kawaguchi. 2019 · 2019
Closest in time.