Fetching the paper…
Reading the bibliography…
Deep learning (DL) creates impactful advances following a virtuous recipe: model architecture search, creating large training data sets, and scaling computation.
A Generalization of the Glivenko-Cantelli Theorem
H. G. Tucker · 1959
Earlier work this paper cites.
A General Lower Bound on the Number of Examples Needed for Learning
A. Ehrenfeucht, D. Haussler, M. Kearns, and L. Valiant · 1988
Earlier work this paper cites.
Quantifying Inductive Bias: AI Learning Algorithms and Valiant’s Learning Framework
D. Haussler · 1988
Earlier work this paper cites.
Learnability and the Vapnik-Chervonenkis Dimension
A. Blumer, A. Ehrenfeucht, D. Haussler, and M. K. Warmuth · 1989
Earlier work this paper cites.
Statistical Theory of Learning a Rule
G. Györgyi and N. Tishby · 1990
Earlier work this paper cites.
Four Types of Learning Curves
S. Amari, N. Fujita, and S. Shinomoto · 1992
Earlier work this paper cites.
Statistical Mechanics of Learning from Examples
H. S. Seung, H. Sompolinsky, and N. Tishby · 1992
Earlier work this paper cites.
A Universal Theorem on Learning Curves
S. Amari · 1993
Earlier work this paper cites.
Statistical Theory of Learning Curves under Entropic Loss Criterion
S. Amari and N. Murata · 1993
Earlier work this paper cites.
Statistical Mechanics of Learning in a Large Committee Machine
H. Schwarze and J. Hertz · 1993
Earlier work this paper cites.
Rigorous Learning Curve Bounds from Statistical Mechanics
D. Haussler, M. Kearns, H. S. Seung, and N. Tishby · 1996
Earlier work this paper cites.
An Overview of Statistical Learning Theory
V. Vapnik · 1998
Earlier work this paper cites.
Scaling to Very Very Large Corpora for Natural Language Disambiguation
M. Banko and E. Brill · 2001
Earlier work this paper cites.
Rademacher and Gaussian Complexities: Risk Bounds and Structural Results
P. L. Bartlett and S. Mendelson · 2002
Cited alongside, same era.
Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Cited alongside, same era.
Moses: Open Source Toolkit for Statistical Machine Translation
P. Koehn, H. Hoang, A. Birch, C. Callison-Burch, M. Federico, N. Bertoldi, B. Cowan, W. Shen, C. Moran, R. Zens, C. Dyer, O. Bojar, A. Constantin, and E. Herbst · 2007
Cited alongside, same era.
One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson · 2013
Cited alongside, same era.
Deep Speech: Scaling Up End-to-End Speech Recognition
A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, et al · 2014
A Closer Look at Memorization in Deep Networks
D. Arpit, S. Jastrzębski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. Courville, Y. Bengio, and S. Lacoste-Julien · 2017
Closest in time.
Exploring Neural Transducers for End-to-end Speech Recognition
E. Battenberg, J. Chen, R. Child, A. Coates, Y. Gaur, Y. Li, H. Liu, S. Satheesh, D. Seetapun, A. Sriram, and Z. Zhu · 2017
Closest in time.
Capacity and Trainability in Recurrent Neural Networks
J. Collins, J. Sohl-Dickstein, and D. Sussillo · 2017
Closest in time.
Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data
G. K. Dziugaite and D. M. Roy · 2017
Closest in time.
Nearly-tight VC-dimension Bounds for Piecewise Linear Neural Networks
N. Harvey, C. Liaw, and A. Mehrabian · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention-based Models for Speech Recognition
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
T. Luong, H. Pham, and C. D. Manning · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Cited alongside, same era.
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, et al · 2016
Cited alongside, same era.
End-to-end Attention-based Large Vocabulary Speech Recognition
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Exploring the Limits of Language Modeling
R. Jozefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y. Wu · 2016
Cited alongside, same era.
K. Kawaguchi, L. P. Kaelbling, and Y. Bengio · 2017
Closest in time.
Opennmt: Open-source toolkit for neural machine translation
G. Klein, Y. Kim, Y. Deng, J. Senellart, and A. M. Rush · 2017
Closest in time.
Neural Machine Translation (seq2seq) Tutorial
M. Luong, E. Brevdo, and R. Zhao · 2017
Closest in time.
A Bayesian Perspective on Generalization and Stochastic Gradient Descent
S. L. Smith and Q. V. Le · 2017
Closest in time.
Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
C. Sun, A. Shrivastava, S. Singh, and A. Gupta · 2017
Closest in time.
Understanding Deep Learning Requires Rethinking Generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Closest in time.
Recurrent Highway Networks
J. G. Zilly, R. K. Srivastava, J. Koutník, and J. Schmidhuber · 2017
Closest in time.