Fetching the paper…
Reading the bibliography…
Deep learning (DL) research yields accuracy and product improvements from both model architecture changes and scale: larger data sets and models, and more computation.
Prediction and Entropy of Printed English
Claude E. Shannon. 1951 · 1951
Earlier work this paper cites.
Tile Size Selection Using Cache Organization and Data Layout. In ACM SIGPLAN Notices , Vol. 30. ACM, 279–290
Stephanie Coleman and Kathryn S. McKinley. 1995 · 1995
Earlier work this paper cites.
Scaling to Very Very Large Corpora for Natural Language Disambiguation. In Association of Computational Linguistics (ACL)
Michele Banko and Eric Brill. 2001 · 2001
Earlier work this paper cites.
Bandwidth Optimal All-reduce Algorithms for Clusters of Workstations
Pitch Patarasuk and Xin Yuan. 2009 · 2009
Earlier work this paper cites.
Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures
Samuel Williams, Andrew Waterman, and David Patterson. 2009 · 2009
Earlier work this paper cites.
Hogwild: A Lock-free Approach to Parallelizing Stochastic Gradient Descent. In Advances in Neural Information Processing Systems (NIPS) . 693–701
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu. 2011 · 2011
Earlier work this paper cites.
Deep Learning with COTS HPC Systems. In International Conference on Machine Learning (ICML) . 1337–1345
Adam Coates, Brody Huval, Tao Wang, David Wu, Bryan Catanzaro, and Andrew Ng. 2013 · 2013
Earlier work this paper cites.
cuDNN: Efficient Primitives for Deep Learning
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014 · 2014
Earlier work this paper cites.
Deep Speech: Scaling Up End-to-End Speech Recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al · 2014
Earlier work this paper cites.
One Weird Trick for Parallelizing Convolutional Neural Networks
Alex Krizhevsky. 2014 · 2014
Earlier work this paper cites.
Hasim Sak, Andrew W. Senior, and Françoise Beaufays. 2014 · 2014
Earlier work this paper cites.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2015 · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation. In The Conference on Empirical Methods in Natural Language Processing (EMNLP) . 1412–1421
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin. In The International Conference on Machine Learning (ICML) . 173–182
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, JingDong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, et al · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Exploring the Limits of Language Modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Cited alongside, same era.
A Bayesian Perspective on Generalization and Stochastic Gradient Descent
Samuel L. Smith and Quoc V. Le. 2017 · 2017
Later among the works it cites.
Revisiting Unreasonable Effectiveness of Data in Deep Learning Era. In The International Conference on Computer Vision (ICCV)
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. 2017 · 2017
Later among the works it cites.
TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning. In Advances in Neural Information Processing Systems (NIPS)
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2017 · 2017
Later among the works it cites.
The Microsoft 2017 Conversational Speech Recognition System
Wayne Xiong, Lingfeng Wu, Fil Alleva, Jasha Droppo, Xuedong Huang, and Andreas Stolcke. 2017 · 2017
Later among the works it cites.
Recurrent Highway Networks. In The International Conference on Machine Learning (ICML)
Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding. In Advances in Neural Information Processing Systems (NIPS) . 1709–1720
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic. 2017 · 2017
Cited alongside, same era.
Exploring Neural Transducers for End-to-end Speech Recognition. In IEEE Automatic Speech Recognition and Understanding Workshop . 206–213
Eric Battenberg, Jitong Chen, Rewon Child, Adam Coates, Yashesh Gaur, Yi Li, Hairong Liu, Sanjeev Satheesh, David Seetapun, Anuroop Sriram, and Zhenyao Zhu. 2017 · 2017
Cited alongside, same era.
Machine Learning for Systems and Systems for Machine Learning
Jeff Dean et al · 2017
Cited alongside, same era.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Priya Goyal, Piotr Dollàr, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Cited alongside, same era.
Nearly-tight VC-dimension and Pseudodimension Bounds for Piecewise Linear Neural Networks. In The Conference on Learning Theory (COLT) , Vol. 65. 1064–1068
Nick Harvey, Peter L. Bartlett, Christopher Liaw, and Abbas Mehrabian. 2017 · 2017
Cited alongside, same era.
Deep Learning Scaling is Predictable, Empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. 2017 · 2017
Cited alongside, same era.
Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training. In The International Conference on Learning Representations (ICLR)
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J. Dally. 2017 · 2017
Cited alongside, same era.
Paleo: A Performance Model for Deep Neural Networks. In The International Conference on Learning Representations (ICLR)
Hang Qi, Evan R. Sparks, and Ameet Talwalkar. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
AI and Compute
2018 · 2018
Later among the works it cites.
DAWNBench
2018 · 2018
Later among the works it cites.
DeepBench
2018 · 2018
Later among the works it cites.
Model Compression and Acceleration for Deep Neural Networks: The Principles, Progress, and Challenges. In IEEE Signal Processing Magazine , Vol. 35
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang. 2018 · 2018
Later among the works it cites.
Semantics-Preserving Parallelization of Stochastic Gradient Descent. In IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 224–233
Saeed Maleki, Madanlal Musuvathi, and Todd Mytkowicz. 2018 · 2018
Later among the works it cites.
MLPerf: A Broad ML Benchmark Suite for Measuring Performance of ML Software Frameworks, ML Hardware Accelerators, and ML Cloud Platforms
MLPerf. 2018 · 2018
Later among the works it cites.
Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
Raul Puri, Robert Kirby, Nikolai Yakovenko, and Bryan Catanzaro. 2018 · 2018
Later among the works it cites.
ImageNet Training in Minutes. In The International Conference on Parallel Processing (ICPP) . 1:1–1:10
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer. 2018 · 2018
Later among the works it cites.