Fetching the paper…
Reading the bibliography…
The generalization error of deep neural networks via their classification margin is studied in this work.
G. Cybenko, “Approximation by superpositions of a sigmoidal function,” Mathematics of Control Signals and Systems , vol. 2, no. 4, pp. 303–314, 1989
1989
Earlier work this paper cites.
K. Hornik, “Approximation capabilities of multilayer feedforward networks,” Neural Networks , vol. 4, no. 2, pp. 251–257, 1991
1991
Earlier work this paper cites.
G. Watson, “Characterization of the subdifferential of some matrix norms,” Linear Algebra and its Applications , vol. 170, pp. 33–45, Jun. 1992
1992
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, Nov. 1997
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, Nov. 1998
1998
Earlier work this paper cites.
V. N. Vapnik, “An overview of statistical learning theory,” IEEE Transactions on Neural Networks , vol. 10, no. 5, pp. 988–999, Sep. 1999
1999
Earlier work this paper cites.
P. L. Bartlett and S. Mendelson, “Rademacher and Gaussian complexities: risk bounds and structural results,” Journal of Machine Learning Research (JMLR) , vol. 3, pp. 463–482, 2002
2002
Earlier work this paper cites.
D. Shi, D. S. Yeung, and J. Gao, “Sensitivity analysis applied to the construction of radial basis function networks,” Neural networks , vol. 18, no. 7, pp. 951–957, Mar. 2005
2005
Earlier work this paper cites.
K. Q. Shen, C. J. Ong, X. P. Li, and E. P. V. Wilder-Smith, “Feature selection via sensitivity analysis of SVM probabilistic outputs,” Machine Learning , vol. 70, no. 1, pp. 1–20, Jan. 2008
2008
Earlier work this paper cites.
J. B. Yang, K. Q. Shen, C. J. Ong, and X. P. Li, “Feature selection via sensitivity analysis of MLP probabilistic outputs,” IEEE International Conference on Systems, Man and Cybernetics , pp. 774–779, 2008
2008
Earlier work this paper cites.
S. Mendelson, A. Pajor, and N. Tomczak-Jaegermann, “Uniform uncertainty principle for Bernoulli and subgaussian ensembles,” Constructive Approximation , vol. 28, no. 3, pp. 277–289, Dec. 2008
2008
Earlier work this paper cites.
Y. Bengio, “Learning deep architectures for AI,” Foundations and trends® in Machine Learning , vol. 2, no. 1, pp. 1–127, 2009
2009
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” Computer Science Department, University of Toronto, Tech. Rep , Apr. 2009
2009
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted Boltzmann machines,” Proceedings of the 27th International Conference on Machine Learning (ICML) , pp. 807–814, 2010
2010
Earlier work this paper cites.
Y.-L. Boureau, J. Ponce, and Y. LeCun, “A theoretical analysis of feature pooling in visual recognition,” Proceedings of the 27th International Conference on Machine Learning (ICML) , pp. 111–118, 2010
2010
Earlier work this paper cites.
S. Rifai, P. Vincent, X. Muller, X. Glorot, and Y. Bengio, “Contractive auto-encoders: explicit invariance during feature extraction,” Proceedings of the 28th International Conference on Machine Learning (ICML) , pp. 833–840, 2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in Neural Information Processing Systems (NIPS) , pp. 1097–1105, 2012
2012
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: the shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, Oct. 2012
2012
Earlier work this paper cites.
S. Mallat, “Group invariant scattering,” Communications on Pure and Applied Mathematics , vol. 65, no. 10, pp. 1331–1398, 2012
2012
Earlier work this paper cites.
J. Bruna and S. Mallat, “Invariant scattering convolution networks,” IEEE Transactions on Pattern Analysis and Machine Intellignce , vol. 35, no. 8, pp. 1872–1886, Mar. 2012
2012
Cited alongside, same era.
H. Xu and S. Mannor, “Robustness and generalization,” Machine Learning , vol. 86, no. 3, pp. 391–423, 2012
2012
Cited alongside, same era.
K. B. Petersen and M. S. Pedersen, “The matrix cookbook,” Technical University of Denmark , 2012
2012
Cited alongside, same era.
J. Bruna, A. Szlam, and Y. LeCun, “Learning stable group invariant representations with convolutional networks,” International Conference on Learning Representations (ICLR) , 2013
2013
Cited alongside, same era.
N. Verma, “Distance preserving embeddings for general n-dimensional manifolds.” Journal of Machine Learning Research (JMLR) , vol. 14, no. 1, pp. 2415–2448, Aug. 2013
2013
J. Huang, Q. Qiu, G. Sapiro, and R. Calderbank, “Discriminative robust transformation learning,” Advances in Neural Information Processing Systems (NIPS) , pp. 1333–1341, 2015
2015
Later among the works it cites.
B. Neyshabur, R. Tomioka, and N. Srebro, “Norm-based capacity control in neural networks,” Proceedings of The 28th Conference on Learning Theory (COLT) , pp. 1376–1401, 2015
2015
Later among the works it cites.
2015
Later among the works it cites.
B. Neyshabur, R. Tomioka, R. Salakhutdinov, and N. Srebro, “Data-dependent path normalization in neural networks,” International Conference on Learning Representations (ICLR) , 2015
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
G. Montúfar, R. Pascanu, K. Cho, and Y. Bengio, “On the number of linear regions of deep neural networks,” Advances in Neural Information Processing Systems (NIPS) , pp. 2924–2932, 2014
2014
Cited alongside, same era.
A. M. Saxe, J. L. McClelland, and S. Ganguli, “Exact solutions to the nonlinear dynamics of learning in deep linear neural networks,” International Conference on Learning Representations (ICLR) , 2014
2014
Cited alongside, same era.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research (JMLR) , vol. 15, no. 1, pp. 1929–1958, Jun. 2014
2014
Cited alongside, same era.
S. Shalev-Shwartz and S. Ben-David, Understanding machine learning: from theory to algorithms . Cambridge University Press, 2014
2014
Cited alongside, same era.
Y. S. Hsiao, J. Sanchez-Riera, T. Lim, K. L. Hua, and W. H. Cheng, “LaRED: a large RGB-D extensible hand gesture dataset,” Proceedings of the 5th ACM Multimedia Systems Conference , pp. 53–58, 2014
2014
Cited alongside, same era.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, May 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet large scale visual recognition challenge,” International Journal of Computer Vision (IJCV) , vol. 115, no. 3, pp. 211–252, 2015
2015
Later among the works it cites.
J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller, “Striving for simplicity: the all convolutional net,” International Conference on Learning Representations (ICLR - workshop track) , Dec. 2015
2015
Later among the works it cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , Dec. 2016
2016
Closest in time.
N. Cohen, O. Sharir, and A. Shashua, “On the expressive power of deep learning: a tensor analysis,” 29th Annual Conference on Learning Theory (COLT) , pp. 698–728, 2016
2016
Closest in time.
M. Telgarsky, “Benefits of depth in neural networks,” 29th Annual Conference on Learning Theory (COLT) , pp. 1517–1539, 2016
2016
Closest in time.
R. Giryes, G. Sapiro, and A. M. Bronstein, “Deep neural networks with random Gaussian weights: a universal classification strategy?” IEEE Transactions on Signal Processing , vol. 64, no. 13, pp. 3444–3457, Jul. 2016
2016
Closest in time.
2016
Closest in time.
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” Advances in Neural Information Processing Systems (NIPS) , pp. 901–909, 2016
2016
Closest in time.
S. Zagoruyko and N. Komodakis, “Wide residual networks,” arXiv:1605.07146 , 2016
2016
Closest in time.
2016
Closest in time.
2016
Closest in time.
2016
Closest in time.
2016
Closest in time.
J. Sokolić, R. Giryes, G. Sapiro, and M. R. D. Rodrigues, “Generalization error of invariant classifiers,” in International Conference on Artificial Intelligence and Statistics (AISTATS) , 2017
2017
Closest in time.