Fetching the paper…
Reading the bibliography…
The wide adoption of DNNs has given birth to unrelenting computing requirements, forcing datacenter operators to adopt domain-specific accelerators to train them.
G. Marsaglia, “Xorshift RNGs,” Journal of Statistical Software
2003
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” Technical report, University of Toronto
2009
Earlier work this paper cites.
T. Mikolov, M. Karafiát, L. Burget, J. Cernocký, and S. Khudanpur, “Recurrent neural network based language model,” in Proceedings of 11th Annual Conference of the International Speech Communication Association (INTERSPEECH 2010)
2010
Earlier work this paper cites.
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng, “Reading digits in natural images with unsupervised feature learning,” Deep Learning and Unsupervised Feature Learning Workshop
2011
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Courbariaux, Y. Bengio, and J.-P. David, “BinaryConnect: Training Deep Neural Networks with binary weights during propagations.,” in Proceedings of Twenty-ninth Conference on Neural Information Processing Systems (NIPS 15)
2015
Earlier work this paper cites.
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan, “Deep Learning with Limited Numerical Precision.,” in Proceedings of Thirty-second International Conference on Machine Learning (ICML 15)
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Binarized Neural Networks.,” in Proceedings of Thirtieth Conference on Neural Information Processing Systems (NIPS 16)
2016
Earlier work this paper cites.
F. Li and B. Liu, “Ternary Weight Networks.,” CoRR
2016
Earlier work this paper cites.
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks.,” in Proceedings of 13rd European Conference on Computer Vision (ECCV 16)
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition.,” in Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR 16)
2016
Cited alongside, same era.
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger, “Deep networks with stochastic depth,” in Proceedings of 13rd European Conference on Computer Vision (ECCV 16)
2016
Cited alongside, same era.
S. Zagoruyko and N. Komodakis, “Wide Residual Networks.,” in Proceedings of British Machine Vision Conference (BMVC 16)
2016
Cited alongside, same era.
Y.-H. Chen, J. S. Emer, and V. Sze, “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks.,” in Proceedings of The 43rd International Symposium on Computer Architecture (ISCA 16)
2016
Cited alongside, same era.
G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger, “Densely Connected Convolutional Networks.,” in Proceedings of Conference on Computer Vision and Pattern Recognition (CVPR 17)
2017
Later among the works it cites.
2017
Later among the works it cites.
Accessed: 2018-01-31
“Artificial intelligence architecture.” https://www.nvidia.com/en-us/data-center/volta-gpu-architecture , 2018 · 2018
Closest in time.
Accessed: 2018-05-15
“Tearing apart google’s tpu 3.0 ai coprocessor.” https://www.nextplatform.com/2018/05/10/tearing-apart-googles-tpu-3-0-ai-coprocessor , 2018 · 2018
Closest in time.
Accessed: 2018-01-31
W. Dally, “High performance hardware for machine learning.” https://media.nips.cc/Conferences/2015/tutorialslides/Dally-NIPS-Tutorial-2015.pdf , 2015 · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
U. Köster, T. Webb, X. Wang, M. Nassar, A. K. Bansal, W. Constable, O. Elibol, S. Hall, L. Hornof, A. Khosrowshahi, C. Kloss, R. J. Pai, and N. Rao, “Flexpoint: An Adaptive Numerical Format for Efficient Training of Deep Neural Networks.,” in Proceedings of Thirty-first Conference on Neural Information Processing Systems (NIPS 17)
2017
Cited alongside, same era.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P. luc Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon, “In-Datacenter Performance Analysis of a Tensor Processing Unit.,” in Proceedings of The 44th International Symposium on Computer Architecture (ISCA 17)
2017
Cited alongside, same era.
C. Zhu, S. Han, H. Mao, and W. J. Dally, “Trained Ternary Quantization.,” in Proceedings of Fifth International Conference on Learning Representations (ICLR 17)
2017
Cited alongside, same era.
H. Zhang, J. Li, K. Kara, D. Alistarh, J. Liu, and C. Zhang, “ZipML: Training Linear Models with End-to-End Low Precision, and a Little Bit of Deep Learning.,” in Proceedings of Thirty-fourth International Conference on Machine Learning (ICML 17)
2017
Cited alongside, same era.
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic, “QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding.,” in Proceedings of Thirty-first Conference on Neural Information Processing Systems (NIPS 17)
2017
Cited alongside, same era.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in pytorch,” NIPS 2017 Autodiff Workshop: The Future of Gradient-based Machine Learning Software and Techniques
2017
Cited alongside, same era.
P. Micikevicius, S. Narang, J. Alben, G. F. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu, “Mixed precision training,” in Proceedings of Sixth International Conference on Learning Representations (ICLR 18)
2018
Closest in time.
Accessed: 2018-01-31
“How to quantize neural networks with tensorflow.” https://www.tensorflow.org/performance/quantization , 2017 · 2018
Closest in time.
Z. Song, Z. Liu, and D. Wang, “Computation error analysis of block floating point arithmetic oriented convolution neural network accelerator design,” in Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18)
2018
Closest in time.
Accessed: 2018-01-31
“Cloud tpu.” https://cloud.google.com/tpu , 2017 · 2018
Closest in time.
S. Merity, N. S. Keskar, and R. Socher, “Regularizing and optimizing LSTM language models,” in Proceedings of Sixth International Conference on Learning Representations (ICLR 18)
2018
Closest in time.