Fetching the paper…
Reading the bibliography…
Training neural networks involves solving large-scale non-convex optimization problems.
Learning internal representations by error propagation
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Marcus, Mitchell P., Santorini, Beatrice, and Marcinkiewicz, Mary Ann · 1993
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Leon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Jarrett, Kevin, Kavukcuoglu, Koray, Ranzato, Marc’Aurelio, and LeCun, Yann · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, Alex and Hinton, Geoffrey · 2009
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
Bergstra, James, Breuleux, Olivier, Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Desjardins, Guillaume, Turian, Joseph, Warde-Farley, David, and Bengio, Yoshua · 2010
Cited alongside, same era.
Convolutional networks and applications in vision
LeCun, Yann, Kavukcuoglu, Koray, and Farabet, Clément · 2010
Cited alongside, same era.
Deep sparse rectifier neural networks
Glorot, Xavier, Bordes, Antoine, and Bengio, Yoshua · 2011
Cited alongside, same era.
Theano: new features and speed improvements
Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Bergstra, James, Goodfellow, Ian J., Bergeron, Arnaud, Bouchard, Nicolas, and Bengio, Yoshua · 2012
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, Andrew M., McClelland, James L., and Ganguli, Surya · 2013
Cited alongside, same era.
Improving Neural Networks with Dropout
Srivastava, Nitish · 2013
The Loss Surface of Multilayer Networks
Choromanska, A., Henaff, M., Mathieu, M., Ben Arous, G., and LeCun, Y · 2014
Closest in time.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Yann N, Pascanu, Razvan, Gulcehre, Caglar, Cho, Kyunghyun, Ganguli, Surya, and Bengio, Yoshua · 2014
Closest in time.
Explaining and harnessing adversarial examples
Goodfellow, Ian, Shlens, Jonathon, and Szegedy, Christian · 2014
Closest in time.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, Nitish, Hinton, Geoffrey, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan · 2014
Closest in time.
Recurrent neural network regularization
Zaremba, Wojciech, Sutskever, Ilya, and Vinyals, Oriol · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-prediction deep Boltzmann machines
Goodfellow, Ian J., Mirza, Mehdi, Courville, Aaron, and Bengio, Yoshua
Cited in the paper.
Pylearn2: a machine learning research library
Goodfellow, Ian J., Warde-Farley, David, Lamblin, Pascal, Dumoulin, Vincent, Mirza, Mehdi, Pascanu, Razvan, Bergstra, James, Bastien, Frédéric, and Bengio, Yoshua
Cited in the paper.
Maxout networks
Goodfellow, Ian J., Warde-Farley, David, Mirza, Mehdi, Courville, Aaron, and Bengio, Yoshua
Cited in the paper.