Fetching the paper…
Reading the bibliography…
Recent work has shown how to train Convolutional Neural Networks (CNNs) rapidly on large image datasets, then transfer the knowledge gained from these models to a variety of tasks.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” 2009
2009
Earlier work this paper cites.
J. Turian, L. Ratinov, and Y. Bengio, “Word representations: A simple and general method for semi-supervised learning,” in Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics , ser. ACL ’10. Stroudsburg, PA, USA: Association for Computational Linguistics, 2010, pp. 384–394
2010
Earlier work this paper cites.
R. Parker, D. Graff, J. Kong, K. Chen, and K. Maeda, “English gigaword fifth edition,” 2011. [Online]. Available: https://catalog.ldc.upenn.edu/ldc2011t07
2011
Earlier work this paper cites.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research , vol. 12, pp. 2825–2830, 2011
2011
Earlier work this paper cites.
Y. A. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller, Efficient BackProp . Berlin, Heidelberg: Springer Berlin Heidelberg, 2012, pp. 9–48. [Online]. Available: https://doi.org/10.1007/978-3-642-35289-8_3
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell, “Decaf: A deep convolutional activation feature for generic visual recognition,” International Conference on Machine Learning , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
I. Sutskever, “Training recurrent neural networks,” 2013. [Online]. Available: https://www.cs.utoronto.ca/~ilya/pubs/ilya_sutskever_phd_thesis.pdf
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
A. S. Razavian, H. Azizpour, J. Sullivan, and S. Carlsson, “CNN features off-the-shelf: an astounding baseline for recognition,” IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , 2014
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” 2014. [Online]. Available: https://www.aclweb.org/anthology/D14-1162
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. McAuley, C. Targett, Q. Shi, and A. van den Hengel, “Image-based recommendations on styles and substitutes,” SIGIR , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. M. Dai and Q. V. Le, “Semi-supervised sequence learning,” CoRR , vol. abs/1511.01432, 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
S. Gray, A. Radford, and D. P. Kingma, “Gpu kernels for block-sparse weights,” 2017. [Online]. Available: https://blog.openai.com/block-sparse-gpu-kernels/
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
E. Hoffer, I. Hubara, and D. Soudry, “Train longer, generalize better: closing the generalization gap in large batch training of neural networks,” 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
O. Melamud, J. Goldberger, and I. Dagan, “context2vec: Learning generic context embedding with bidirectional lstm,” in Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning , 01 2016, pp. 51–61
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
K. He, G. Gkioxari, P. Dollár, and R. B. Girshick, “Mask R-CNN,” IEEE International Conference on Computer Vision (ICCV) , 2017
2017
Cited alongside, same era.
2017
Later among the works it cites.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in pytorch,” 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
S. Kornblith, J. Shlens, and Q. V. Le, “Do better imagenet models transfer better?” 2018
2018
Closest in time.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” 2018. [Online]. Available: https://blog.openai.com/language-unsupervised/
2018
Closest in time.
M. Ott, S. Edunov, D. Grangier, and M. Auli, “Scaling neural machine translation,” 2018
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
P. J. Liu, M. Saleh, E. Pot, B. Goodrich, R. Sepassi, L. Kaiser, and N. Shazeer, “Generating wikipedia by summarizing long sequences,” ICLR , 2018
2018
Closest in time.
NVIDIA. (2018) Mixed precision training: Choosing a scaling factor. [Online]. Available: https://docs.nvidia.com/deeplearning/sdk/mixed-precision-training/index.html#scalefactor
2018
Closest in time.