Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Incremental gradient, subgradient, and proximal methods for convex optimization: A survey
Dimitri P Bertsekas · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
Beneath the valley of the noncommutative arithmetic-geometric mean inequality: conjectures, case-studies
Benjamin Recht and Christopher Re · 2012
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Original
Ross B. Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Earlier work this paper cites.
Deepface: Closing the gap to human-level performance in face verification
Yaniv Taigman, Ming Yang, Marc’Aurelio Ranzato, and Lior Wolf · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
1.1 computing’s energy problem (and what we can do about it)
Mark Horowitz · 2014
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Why random reshuffling beats stochastic gradient descent
Original
Mert Gürbüzbalaban, Asu Ozdaglar, and Pablo Parrilo · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Variation-tolerant architectures for convolutional neural networks in the near threshold voltage regime
Y. Lin, S. Zhang, and N. R. Shanbhag · 2016
Earlier work this paper cites.
Qsgd: Randomized quantization for communication-optimal stochastic gradient descent
Original
Dan Alistarh, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2016
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J Dally · 2016
Earlier work this paper cites.
Without-replacement sampling for stochastic gradient methods
Ohad Shamir · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Original
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks, 2016
Andreas Veit, Michael Wilber, and Serge Belongie · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger · 2016
Earlier work this paper cites.
Designing energy-efficient convolutional neural networks using energy-aware pruning
Tien-Ju Yang, Yu-Hsin Chen, and Vivienne Sze · 2017
Earlier work this paper cites.
One model to learn them all
Original
Lukasz Kaiser, Aidan N. Gomez, Noam Shazeer, Ashish Vaswani, Niki Parmar, Llion Jones, and Jakob Uszkoreit · 2017
Earlier work this paper cites.
Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks
Y. H. Chen, T. Krishna, J. S. Emer, and V. Sze · 2017
Earlier work this paper cites.