Fetching the paper…
Reading the bibliography…
Despite their overwhelming capacity to overfit, deep learning architectures tend to generalize relatively well to unseen data, allowing them to be deployed in practice.
Keeping the neural networks simple by minimizing the description length of the weights
Hinton, Geoffrey E and Van Camp, Drew · 1993
Earlier work this paper cites.
Flat minima
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, Shun-Ichi · 1998
Earlier work this paper cites.
Nonlinear independent component analysis: Existence and uniqueness results
Hyvärinen, Aapo and Pajunen, Petteri · 1999
Earlier work this paper cites.
Random walks on symmetric spaces and inequalities for matrix spectra
Klyachko, Alexander A · 2000
Earlier work this paper cites.
Gaussianization
Chen, Scott Saobing and Gopinath, Ramesh A · 2001
Earlier work this paper cites.
Stability and generalization
Bousquet, Olivier and Elisseeff, André · 2002
Earlier work this paper cites.
On-line learning for very large datasets
Bottou, Léon and LeCun, Yann · 2005
Earlier work this paper cites.
The tradeoffs of large scale learning
Bottou, Léon and Bousquet, Olivier · 2008
Earlier work this paper cites.
Confidence level solutions for stochastic programming
Nesterov, Yurii and Vial, Jean-Philippe · 2008
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Jarrett, Kevin, Kavukcuoglu, Koray, LeCun, Yann, et al · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, Léon · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, Vinod and Hinton, Geoffrey E · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks
Glorot, Xavier, Bordes, Antoine, and Bengio, Yoshua · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Earlier work this paper cites.
Maxout networks
Goodfellow, Ian J, Warde-Farley, David, Mirza, Mehdi, Courville, Aaron C, and Bengio, Yoshua · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Graves, Alex, Mohamed, Abdel-rahman, and Hinton, Geoffrey · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, Andrew M., McClelland, James L., and Ganguli, Surya · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, Kyunghyun, van Merrienboer, Bart, Gülçehre, Çaglar, Bahdanau, Dzmitry, Bougares, Fethi, Schwenk, Holger, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Yann N., Pascanu, Razvan, Gülçehre, Çaglar, Cho, KyungHyun, Ganguli, Surya, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Nice: Non-linear independent components estimation
Dinh, Laurent, Krueger, David, and Bengio, Yoshua · 2014
Cited alongside, same era.
Deep speech: Scaling up end-to-end speech recognition
Hannun, Awni Y., Case, Carl, Casper, Jared, Catanzaro, Bryan, Diamos, Greg, Elsen, Erich, Prenger, Ryan, Satheesh, Sanjeev, Sengupta, Shubho, Coates, Adam, and Ng, Andrew Y · 2014
Cited alongside, same era.
On the number of linear regions of deep neural networks
Montufar, Guido F, Pascanu, Razvan, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Cited alongside, same era.
Revisiting natural gradient for deep networks
Pascanu, Razvan and Bengio, Yoshua · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc V · 2014
Cited alongside, same era.
Optimization methods for large-scale machine learning
Bottou, Léon, Curtis, Frank E, and Nocedal, Jorge · 2016
Later among the works it cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
Chan, William, Jaitly, Navdeep, Le, Quoc V., and Vinyals, Oriol · 2016
Later among the works it cites.
Wav2letter: an end-to-end convnet-based speech recognition system
Collobert, Ronan, Puhrsch, Christian, and Synnaeve, Gabriel · 2016
Later among the works it cites.
Density estimation using real nvp
Dinh, Laurent, Sohl-Dickstein, Jascha, and Bengio, Samy · 2016
Later among the works it cites.
A convolutional encoder model for neural machine translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Szegedy, Christian, Zaremba, Wojciech, Sutskever, Ilya, Bruna, Joan, Erhan, Dumitru, Goodfellow, Ian, and Fergus, Rob · 2014
Cited alongside, same era.
Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015 , volume 37 of JMLR Workshop and Conference Proceedings , 2015. JMLR.org
Bach, Francis R. and Blei, David M. (eds.) · 2015
Cited alongside, same era.
Understanding symmetries in deep networks
Badrinarayanan, Vijay, Mishra, Bamdev, and Cipolla, Roberto · 2015
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Cited alongside, same era.
The loss surfaces of multilayer networks
Choromanska, Anna, Henaff, Mikael, Mathieu, Michaël, Arous, Gérard Ben, and LeCun, Yann · 2015
Cited alongside, same era.
Attention-based models for speech recognition
Chorowski, Jan K, Bahdanau, Dzmitry, Serdyuk, Dmitriy, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Cited alongside, same era.
Natural neural networks
Desjardins, Guillaume, Simonyan, Karen, Pascanu, Razvan, and Kavukcuoglu, Koray · 2015
Cited alongside, same era.
Gehring, Jonas, Auli, Michael, Grangier, David, and Dauphin, Yann N · 2016
Later among the works it cites.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, Moritz, Recht, Ben, and Singer, Yoram · 2016
Later among the works it cites.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2016
Later among the works it cites.
An empirical analysis of deep network loss surfaces
Im, Daniel Jiwoong, Tao, Michael, and Branson, Kristin · 2016
Later among the works it cites.
Improved variational inference with inverse autoregressive flow
Kingma, Diederik P, Salimans, Tim, Jozefowicz, Rafal, Chen, Xi, Sutskever, Ilya, and Welling, Max · 2016
Later among the works it cites.
About diagonal rescaling applied to neural nets
Lafond, Jean, Vasilache, Nicolas, and Bottou, Léon · 2016
Later among the works it cites.
On the expressive power of deep neural networks
Raghu, Maithra, Poole, Ben, Kleinberg, Jon, Ganguli, Surya, and Sohl-Dickstein, Jascha · 2016
Later among the works it cites.
Singularity of the hessian in deep learning
Sagun, Levent, Bottou, Léon, and LeCun, Yann · 2016
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, Tim and Kingma, Diederik P · 2016
Later among the works it cites.
Improved techniques for training gans
Salimans, Tim, Goodfellow, Ian, Zaremba, Wojciech, Cheung, Vicki, Radford, Alec, and Chen, Xi · 2016
Later among the works it cites.
Local minima in training of deep networks
Swirszcz, Grzegorz, Czarnecki, Wojciech Marian, and Pascanu, Razvan · 2016
Later among the works it cites.
A note on the evaluation of generative models
Theis, Lucas, Oord, Aäron van den, and Bethge, Matthias · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Yonghui, Schuster, Mike, Chen, Zhifeng, Le, Quoc V, Norouzi, Mohammad, Macherey, Wolfgang, Krikun, Maxim, Cao, Yuan, Gao, Qin, Macherey, Klaus, et al · 2016
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, Pratik, Choromanska, Anna, Soatto, Stefano, LeCun, Yann, Baldassi, Carlo, Borgs, Christian, Chayes, Jennifer, Sagun, Levent, and Zecchina, Riccardo · 2017
Closest in time.
Fast rates for empirical risk minimization of strict saddle problems
Gonen, Alon and Shalev-Shwartz, Shai · 2017
Closest in time.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, Nitish Shirish, Mudigere, Dheevatsa, Nocedal, Jorge, Smelyanskiy, Mikhail, and Tang, Ping Tak Peter · 2017
Closest in time.
Understanding deep learning requires rethinking generalization
Zhang, Chiyuan, Bengio, Samy, Hardt, Moritz, Recht, Benjamin, and Vinyals, Oriol · 2017
Closest in time.