Fetching the paper…
Reading the bibliography…
We examine the role of memorization in deep learning, drawing connections to capacity, generalization, and adversarial robustness.
Variabilita e mutabilita
Gini, Corrado · 1913
Earlier work this paper cites.
Discriminatory analysis-nonparametric discrimination: consistency properties
Fix, Evelyn and Hodges Jr, Joseph L · 1951
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, George · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, Kurt, Stinchcombe, Maxwell, and White, Halbert · 1989
Earlier work this paper cites.
Training with noise is equivalent to tikhonov regularization
Bishop, Chris M · 1995
Earlier work this paper cites.
Overtraining, regularization and searching for a minimum, with application to neural networks
Sjoberg, J., Sjoeberg, J., Sjöberg, J., and Ljung, L · 1995
Earlier work this paper cites.
The effects of adding noise during backpropagation training on a generalization performance
An, Guozhong · 1996
Earlier work this paper cites.
Online learning and stochastic approximations
Bottou, Léon · 1998
Earlier work this paper cites.
The mnist database of handwritten digits, 1998
LeCun, Yann, Cortes, Corinna, and Burges, Christopher JC · 1998
Earlier work this paper cites.
Statistical learning theory , volume 1
Vapnik, Vladimir Naumovich and Vapnik, Vlamimir · 1998
Earlier work this paper cites.
The general inefficiency of batch training for gradient descent learning
Wilson, D Randall and Martinez, Tony R · 2003
Earlier work this paper cites.
Local rademacher complexities
Bartlett, Peter L, Bousquet, Olivier, Mendelson, Shahar, et al · 2005
Earlier work this paper cites.
On early stopping in gradient descent learning
Yao, Yuan, Rosasco, Lorenzo, and Caponnetto, Andrea · 2007
Earlier work this paper cites.
Learning deep architectures for ai
Bengio, Yoshua et al · 2009
Cited alongside, same era.
Kernel analysis of deep networks
Montavon, Grégoire, Braun, Mikio L., and Müller, Klaus-Robert · 2011
Cited alongside, same era.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Goodfellow, Ian J, Mirza, Mehdi, Xiao, Da, Courville, Aaron, and Bengio, Yoshua · 2013
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, Andrew M, McClelland, James L, and Ganguli, Surya · 2013
Cited alongside, same era.
Intriguing properties of neural networks
Szegedy, Christian, Zaremba, Wojciech, Sutskever, Ilya, Bruna, Joan, Erhan, Dumitru, Goodfellow, Ian J., and Fergus, Rob · 2013
Cited alongside, same era.
Deep Learning
Goodfellow, Ian, Bengio, Yoshua, and Courville, Aaron · 2016
Later among the works it cites.
An empirical analysis of deep network loss surfaces
Im, Daniel Jiwoong, Tao, Michael, and Branson, Kristin · 2016
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, Nitish Shirish, Mudigere, Dheevatsa, Nocedal, Jorge, Smelyanskiy, Mikhail, and Tang, Ping Tak Peter · 2016
Later among the works it cites.
Adversarial examples in the physical world
Kurakin, Alexey, Goodfellow, Ian, and Bengio, Samy · 2016
Later among the works it cites.
Why does deep and cheap learning work so well?
Lin, Henry W and Tegmark, Max · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Explaining and harnessing adversarial examples
Goodfellow, Ian J, Shlens, Jonathon, and Szegedy, Christian · 2014
Cited alongside, same era.
On the number of linear regions of deep neural networks
Montufar, Guido F, Pascanu, Razvan, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, Behnam, Tomioka, Ryota, and Srebro, Nathan · 2014
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, Moritz, Recht, Benjamin, and Singer, Yoram · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, Dougal, Duvenaud, David K, and Adams, Ryan P · 2015
Cited alongside, same era.
Distributional smoothing with virtual adversarial training
Miyato, Takeru, Maeda, Shin-ichi, Koyama, Masanori, Nakae, Ken, and Ishii, Shin · 2015
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, Pratik, Choromanska, Anna, Soatto, Stefano, and LeCun, Yann · 2016
Cited alongside, same era.
Later among the works it cites.
Exponential expressivity in deep neural networks through transient chaos
Poole, Ben, Lahiri, Subhaneil, Raghu, Maithreyi, Sohl-Dickstein, Jascha, and Ganguli, Surya · 2016
Later among the works it cites.
On the expressive power of deep neural networks
Raghu, Maithra, Poole, Ben, Kleinberg, Jon, Ganguli, Surya, and Sohl-Dickstein, Jascha · 2016
Later among the works it cites.
Robust large margin deep neural networks
Sokolic, Jure, Giryes, Raja, Sapiro, Guillermo, and Rodrigues, Miguel RD · 2016
Later among the works it cites.
Theano: A Python framework for fast computation of mathematical expressions
Theano Development Team, and others · 2016
Later among the works it cites.
Unsupervised Learning by Predicting Noise
Bojanowski, P. and Joulin, A · 2017
Closest in time.
Understanding black-box predictions via influence functions
Koh, Pang Wei and Liang, Percy · 2017
Closest in time.
Understanding deep learning requires rethinking generalization
Zhang, Chiyuan, Bengio, Samy, Hardt, Moritz, Recht, Benjamin, and Vinyals, Oriol · 2017
Closest in time.