Beyond regression: New tools for predictions and analysis in the behavioral science. Cambridge, MA, itd
PJ Werbos · 1974
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
K. Hornik, M. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Back propagation fails to separate where perceptrons succeed
Martin Brady, Raghu Raghavan, and Joseph Slawny · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
Backpropagation converges for multi-layered networks and linearly-separable patterns
M. Gori and A. Tesi · 1991
Earlier work this paper cites.
On the problem of local minima in backpropagation
M. Gori and A. Tesi · 1992
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
G. Hinton and D. Van Camp · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R Barron · 1994
Earlier work this paper cites.
Numerical Optimization
Stephen J Wright and Jorge Nocedal · 1999
Earlier work this paper cites.
An overview of statistical learning theory
Vladimir N Vapnik · 1999
Earlier work this paper cites.
The information bottleneck method
N. Tishby, F. C. Pereira, and W. Bialek · 2000
Earlier work this paper cites.
Boosting algorithms as gradient descent
Llew Mason, Jonathan Baxter, Peter L. Bartlett, and Marcus R. Frean · 2000
Earlier work this paper cites.
Greedy function approximation: a gradient boosting machine
Jerome H Friedman · 2001
Earlier work this paper cites.
Rademacher and Gaussian complexities: risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Vapnik-Chervonenkis dimension of neural nets
P. Bartlett and W. Maass · 2003
Earlier work this paper cites.
Learning over sets using kernel principal angles
L. Wolf and A. Shashua · 2003
Earlier work this paper cites.
Convex neural networks
Yoshua Bengio, Nicolas L Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte · 2005
Earlier work this paper cites.
Near-optimal signal recovery from random projections: Universal encoding strategies?
E. J. Candès and T. Tao · 2006
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, R Socher, Li-Jia Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
On the set of images modulo viewpoint and contrast changes
G. Sundaramoorthi, P. Petersen, V. S. Varadarajan, and S. Soatto · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G.E. Hinton · 2010
Earlier work this paper cites.
Online learning for matrix factorization and sparse coding
Julien Mairal, Francis Bach, Jean Ponce, and Guillermo Sapiro · 2010
Earlier work this paper cites.
Object detection with discriminatively trained part-based models
P. Felzenszwalb, R. Girshick, D. McAllester, and D. Ramanan · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.