Fetching the paper…
Reading the bibliography…
The large amount of online data and vast array of computing resources enable current researchers in both industry and academia to employ the power of deep learning with neural networks.
Multilayer feedforward networks are universal approximators
K. Hornik, M. B. Stinchcombe, and H. White · 1989
Earlier work this paper cites.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Predicting good probabilities with supervised learning
A. Niculescu-Mizil and R. Caruana · 2005
Earlier work this paper cites.
Manifold regularization: A geometric framework for learning from labeled and unlabeled examples
M. Belkin, P. Niyogi, and V. Sindhwani · 2006
Earlier work this paper cites.
On transductive regression
C. Cortes and M. Mohri · 2006
Earlier work this paper cites.
Boosting for transfer learning
W. Dai, Q. Yang, G.-R. Xue, and Y. Yu · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and F.-F. Li · 2009
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
A. Krizhevsky · 2009
Earlier work this paper cites.
A survey on transfer learning
S. J. Pan and Q. Yang · 2010
Earlier work this paper cites.
Caltech-UCSD Birds 200
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
A. Coates, A. Y. Ng, and H. Lee · 2011
Cited alongside, same era.
Domain adaptation in regression
C. Cortes and M. Mohri · 2011
Cited alongside, same era.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng · 2011
Cited alongside, same era.
Algorithms for learning kernels based on centered alignment
C. Cortes, M. Mohri, and A. Rostamizadeh · 2012
Cited alongside, same era.
Low rank approximation and regression in input sparsity time
K. L. Clarkson and D. P. Woodruff · 2013
Cited alongside, same era.
Domain generalization via invariant feature representation
K. Muandet, D. Balduzzi, and B. Schölkopf · 2013
Cited alongside, same era.
Distilling the knowledge in a neural network
G. E. Hinton, O. Vinyals, and J. Dean · 2015
Later among the works it cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F.-F. Li · 2015
Later among the works it cites.
Revisiting the nyström method for improved large-scale machine learning
A. Gittens and M. W. Mahoney · 2016
Later among the works it cites.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Distilling a neural network into a soft decision tree
N. Frosst and G. E. Hinton · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fast and scalable polynomial kernels via explicit feature maps
N. Pham and R. Pagh · 2013
Cited alongside, same era.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Cited alongside, same era.
Sketching as a tool for numerical linear algebra
D. P. Woodruff · 2014
Cited alongside, same era.
How transferable are features in deep neural networks?
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson · 2014
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
On calibration of modern neural networks
C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger · 2017
Later among the works it cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
Deep learning for classical japanese literature
T. Clanuwat, M. Bober-Irizar, A. Kitamoto, A. Lamb, K. Yamamoto, and D. Ha · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Later among the works it cites.
Understanding sparse jl for feature hashing
M. Jagadeesan · 2019
Later among the works it cites.
Why are big data matrices approximately low rank?
M. Udell and A. Townsend · 2019
Later among the works it cites.