Fetching the paper…
Reading the bibliography…
Good old on-line back-propagation for plain multi-layer perceptrons yields a very low 0.35% error rate on the famous MNIST handwritten digits benchmark.
Beyond Regression: New Tools for Prediction and Analysis in the Behavioral Sciences
Werbos, P.J. (1974) · 1974
Earlier work this paper cites.
Une procédure d’apprentissage pour réseau à seuil asymétrique
LeCun, Y. (1985) · 1985
Earlier work this paper cites.
Learning internal representations by error propagation
Rumelhart, D. E., Hinton, G. E. & Williams, R. J. (1986) · 1986
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen
Hochreiter, S. (1991) · 1991
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. & Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Gradient-Based Learning Applied to Document Recognition
LeCun, Y., Bottou, L., Bengio, Y. & Haffner, P. (1998) · 1998
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
Hochreiter, S., Bengio, Y., Frasconi, P. & Schmidhuber, J. (2001) · 2001
Earlier work this paper cites.
Training Invariant Support Vector Machines
Decoste, D. & Scholkopf, B. (2002) · 2002
Cited alongside, same era.
Artificial Intelligence: A Modern Approach (2nd Edition)
Russell, S. & Norvig, P. (2002) · 2002
Cited alongside, same era.
Best Practices for Convolutional Neural Networks Applied to Visual Document Analysis
Simard, P.Y., Steinkraus, D. & Platt, J.C. (2003) · 2003
Cited alongside, same era.
GPUs for Machine Learning Algorithms
Steinkraus, D., Buck, I. & Simard, P.Y. (2005) · 2005
Cited alongside, same era.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D. & Larochelle, H. (2006) · 2006
Cited alongside, same era.
High Performance Convolutional Neural Networks for Document Processing
Chellapilla, K., Puri, S. & Simard, P. (2006) · 2006
Cited alongside, same era.
To recognize shapes, first learn to generate images
Hinton, G. (2007) · 2007
Later among the works it cites.
A trainable feature extractor for handwritten digit recognition
Lauer, F., Suen, C. & Bloch, G. (2007) · 2007
Later among the works it cites.
Unsupervised Learning of Invariant Feature Hierarchies with Applications to Object Recognition
Ranzato, M., Huang, F., Boureau, Y. & LeCun, Y.(2007) · 2007
Later among the works it cites.
Learning a Nonlinear Embedding by Preserving Class Neighborhood Structure
Salakhutdinov, R. & Hinton, G. (2007) · 2007
Later among the works it cites.
NVIDIA CUDA. Reference Manual
NVIDIA. (2009) · 2009
Later among the works it cites.
Optimizing Matrix Transpose in CUDA
Ruetsch, G. & Micikevicius, P. (2009) · 2009
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient Learning of Sparse Representations with an Energy-Based Model
Ranzato, M., Poultney, C., Chopra, S. & LeCun, Y. (2006) · 2006
Cited alongside, same era.
Accelerating Large-scale Convolutional Neural Networks with Parallel Graphics Multiprocessors
Scherer, D. & Behnke, S. (2009) · 2009
Later among the works it cites.