Fetching the paper…
Reading the bibliography…
Researchers have proposed hardware, software, and algorithmic optimizations to improve the computational performance of deep learning.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Roofline: An Insightful Visual Performance Model for Multicore Architectures
Samuel Williams, Andrew Waterman, and David Patterson · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Deep Sparse Rectifier Neural Networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
Hogwild: A Lock-free Approach to Parallelizing Stochastic Gradient Descent
Feng Niu, Benjamin Recht, Christopher Re, and Stephen Wright · 2011
Earlier work this paper cites.
Workload analysis of a large-scale key-value store
Berk Atikoglu, Yuehai Xu, Eitan Frachtenberg, Song Jiang, and Mike Paleczny · 2012
Earlier work this paper cites.
Large Scale Distributed Deep Networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Earlier work this paper cites.
On the Importance of Initialization and Momentum in Deep Learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
cuDNN: Efficient Primitives for Deep Learning
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer · 2014
Earlier work this paper cites.
Project Adam: Building an Efficient and Scalable Deep Learning Training System
Trishul M Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman · 2014
Earlier work this paper cites.
Caffe: Convolutional Architecture for Fast Feature Embedding
Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su · 2014
Earlier work this paper cites.
Fast Large-scale Optimization by Unifying Stochastic Gradient and Quasi-Newton Methods
Jascha Sohl-Dickstein, Ben Poole, and Surya Ganguli · 2014
Earlier work this paper cites.
Dimmwitted: A Study of Main-memory Statistical Analytics
Ce Zhang and Christopher Ré · 2014
Earlier work this paper cites.
Comparative Study of Deep Learning Software Frameworks
Soheil Bahrampour, Naveen Ramakrishnan, Lukas Schott, and Mohak Shah · 2015
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
TensorFlow: A System for Large-Scale Machine Learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Earlier work this paper cites.
Fathom: Reference Workloads for Modern Deep Learning Methods
Robert Adolf, Saketh Rama, Brandon Reagen, Gu-Yeon Wei, and David Brooks · 2016
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Addressing the Straggler Problem for Iterative Convergent Parallel ML
Aaron Harlap, Henggang Cui, Wei Dai, Jinliang Wei, Gregory Ganger, Phillip Gibbons, Garth Gibson, and Eric Xing · 2016
Cited alongside, same era.
SqueezeNet: AlexNet-level Accuracy with 50x Fewer Parameters and
Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Cited alongside, same era.
Asynchrony begets momentum, with an application to deep learning
Ioannis Mitliagkas, Ce Zhang, Stefan Hadjis, and Christopher Ré · 2016
Cited alongside, same era.
Benchmarking of CNNs for Low-Cost, Low-Power Robotics Applications
Dexmont Pena, Andrew Forembski, Xiaofan Xu, and David Moloney · 2017
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, and Quoc V Le · 2017
Later among the works it cites.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
https://mlperf.org/ , 2018
MLPerf · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benchmarking State-of-the-Art Deep Learning Software Tools
Shaohuai Shi, Qiang Wang, Pengfei Xu, and Xiaowen Chu · 2016
Cited alongside, same era.
Second conference on machine translation, 2017
2017
Cited alongside, same era.
https://www.tensorflow.org/performance/xla , 2017
Tensorflow xla overview · 2017
Cited alongside, same era.
Extremely large minibatch sgd: Training resnet-50 on imagenet in 15 minutes
Takuya Akiba, Shuji Suzuki, and Keisuke Fukuda · 2017
Cited alongside, same era.
DeepBench: Benchmarking Deep Learning Operations on Different Hardware
Baidu · 2017
Cited alongside, same era.
Microsoft unveils Project Brainwave for Real-time AI
Doug Burger · 2017
Cited alongside, same era.
Understanding reduced-voltage operation in modern dram devices: Experimental characterization, analysis, and mechanisms
Kevin K Chang, A Giray Yağlıkçı, Saugata Ghose, Aditya Agrawal, Niladrish Chatterjee, Abhijith Kashyap, Donghyuk Lee, Mike O’Connor, Hasan Hassan, and Onur Mutlu · 2017
Cited alongside, same era.
Dario Amodei and Danny Hernandez · 2018
Closest in time.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2018
Closest in time.
What is underfitting and overfitting in machine learning and how to deal with it, 2018
Anup Bhande · 2018
Closest in time.
High-accuracy low-precision training
Christopher De Sa, Megan Leszczynski, Jian Zhang, Alana Marzoev, Christopher R Aberger, Kunle Olukotun, and Christopher Ré · 2018
Closest in time.
What your dram power models are not telling you: Lessons from a detailed experimental study
Saugata Ghose, Abdullah Giray Yaglikçi, Raghav Gupta, Donghyuk Lee, Kais Kudrolli, William X Liu, Hasan Hassan, Kevin K Chang, Niladrish Chatterjee, Aditya Agrawal, et al · 2018
Closest in time.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Closest in time.
Nvidia tensor core programmability, performance & precision
Stefano Markidis, Steven Wei Der Chien, Erwin Laure, Ivy Bo Peng, and Jeffrey S Vetter · 2018
Closest in time.
Revisiting Small Batch Training for Deep Neural Networks
Dominic Masters and Carlo Luschi · 2018
Closest in time.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Closest in time.
Regularized Evolution for Image Classifier Architecture Search
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le · 2018
Closest in time.
Do CIFAR-10 classifiers generalize to cifar-10?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2018
Closest in time.
Horovod: fast and easy distributed deep learning in tensorflow
Alexander Sergeev and Mike Del Balso · 2018
Closest in time.
Imagenet training in minutes
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2018
Closest in time.
Tbd: Benchmarking and analyzing deep neural network training
Hongyu Zhu, Mohamed Akrout, Bojian Zheng, Andrew Pelegris, Amar Phanishayee, Bianca Schroeder, and Gennady Pekhimenko · 2018
Closest in time.
Making ncf reflect production usage, 2019
Victor Bittorf · 2019
Closest in time.
Bigdl: Distributed deep learning library for apache spark, 2019
Intel · 2019
Closest in time.
Beyond data and model parallelism for deep neural networks
Zhihao Jia, Matei Zaharia, and Alex Aiken · 2019
Closest in time.