Modeling by shortest data description
Jorma Rissanen · 1978
Earlier work this paper cites.
Comparison of classifier methods: a case study in handwritten digit recognition
Léon Bottou, Corinna Cortes, John S Denker, Harris Drucker, Isabelle Guyon, Lawrence D Jackel, Yann LeCun, Urs A Muller, Edward Sackinger, Patrice Simard, et al · 1994
Earlier work this paper cites.
Choosing the forcing terms in an inexact newton method
Stanley C Eisenstat and Homer F Walker · 1996
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Neural network learning: Theoretical foundations
Martin Anthony and Peter L Bartlett · 2009
Earlier work this paper cites.
Deep belief networks are compact universal approximators
Nicolas Le Roux and Yoshua Bengio · 2010
Earlier work this paper cites.
Shallow vs. deep sum-product networks
Olivier Delalleau and Yoshua Bengio · 2011
Earlier work this paper cites.
Training deep and recurrent networks with hessian-free optimization
James Martens and Ilya Sutskever · 2012
Earlier work this paper cites.
Intriguing properties of neural networks
Original
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Original
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.