Fetching the paper…
Reading the bibliography…
Background: It is still an open research area to theoretically understand why Deep Neural Networks (DNNs)---equipped with many more parameters than training data and trained by (stochastic) gradient-based methods---often achieve remarkably low generalization error.
Shannon, C. E. (1949), ‘Communication in the presence of noise’, Proceedings of the IRE
1949
Earlier work this paper cites.
Yen, J. (1956), ‘On nonuniform sampling of bandwidth-limited signals’, IRE Transactions on circuit theory
1956
Earlier work this paper cites.
Cybenko, G. (1989), ‘Approximation by superpositions of a sigmoidal function’, Mathematics of control, signals and systems
1989
Earlier work this paper cites.
Dong, D. W. & Atick, J. J. (1995), ‘Statistics of natural time-varying images’, Network: Computation in Neural Systems
1995
Earlier work this paper cites.
Hochreiter, S. & Schmidhuber, J. (1995), Simplifying neural nets by discovering flat minima, in
1995
Earlier work this paper cites.
LeCun, Y. (1998), ‘The mnist database of handwritten digits’, http://yann. lecun. com/exdb/mnist/
1998
Earlier work this paper cites.
Vapnik, V. N. (1999), ‘An overview of statistical learning theory’, IEEE transactions on neural networks
1999
Earlier work this paper cites.
Bousquet, O. & Elisseeff, A. (2002), ‘Stability and generalization’, Journal of machine learning research
2002
Earlier work this paper cites.
Mishali, M. & Eldar, Y. C. (2009), ‘Blind multiband signal reconstruction: Compressed sensing for analog signals’, IEEE Transactions on Signal Processing
2009
Cited alongside, same era.
2014
Cited alongside, same era.
2015
Cited alongside, same era.
LeCun, Y., Bengio, Y. & Hinton, G. (2015), ‘Deep learning’, nature
2015
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
Nagarajan, V. & Kolter, J. Z. (2017), Generalization in deep networks: The role of distance from initialization, in
2017
Later among the works it cites.
2017
Later among the works it cites.
Poggio, T., Kawaguchi, K., Liao, Q., Miranda, B., Rosasco, L., Boix, X., Hidary, J. & Mhaskar, H. (2018), Theory of deep learning iii: the non-overfitting puzzle, Technical report, Technical report, CBMM memo 073
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Closest in time.
2018
Closest in time.