Fetching the paper…
Reading the bibliography…
Recurrent Neural Network (RNN) applications form a major class of AI-powered, low-latency data center workloads.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Dadiannao: A machine-learning supercomputer
Chen, Y., Luo, T., Liu, S., Zhang, S., He, L., Wang, J., Li, L., Chen, T., Xu, Z., Sun, N., et al · 2014
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
Chetlur, S., Woolley, C., Vandermersch, P., Cohen, J., Tran, J., Catanzaro, B., and Shelhamer, E · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Sda: Software-defined accelerator for largescale dnn systems
Ouyang, J., Lin, S., Qi, W., Wang, Y., Yu, B., and Jiang, S · 2014
Earlier work this paper cites.
A reconfigurable fabric for accelerating large-scale datacenter services
Putnam, A., Caulfield, A. M., Chung, E. S., Chiou, D., Constantinides, K., Demme, J., Esmaeilzadeh, H., Fowers, J., Gopal, G. P., Gray, J., Haselman, M., Hauck, S., Heil, S., Hormati, A., Kim, J.-Y., Lanka, S., Larus, J., Peterson, E., Pope, S., Smith, A., Thong, J., Xiao, P. Y., and Burger, D · 2014
Earlier work this paper cites.
Recurrent neural networks hardware implementation on fpga
Chang, A. X. M., Martini, B., and Culurciello, E · 2015
Earlier work this paper cites.
Altera’s 30 billion transistor fpga
Gazettabyte, R. R · 2015
Earlier work this paper cites.
Tensorflow: a system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Earlier work this paper cites.
Eie: efficient inference engine on compressed deep neural network
Han, S., Liu, X., Mao, H., Pu, J., Pedram, A., Horowitz, M. A., and Dally, W. J · 2016
Cited alongside, same era.
Automatic generation of efficient accelerators for reconfigurable hardware
Koeplinger, D., Prabhakar, R., Zhang, Y., Delimitrou, C., Kozyrakis, C., and Olukotun, K · 2016
Cited alongside, same era.
Efficient and reliable high-level synthesis design space explorer for fpgas
Liu, D. and Schafer, B. C · 2016
Cited alongside, same era.
Compression of neural machine translation models via pruning
See, A., Luong, M.-T., and Manning, C. D · 2016
Cited alongside, same era.
Ec2 f1 instances with fpgas now generally available
Amazon · 2017
Cited alongside, same era.
Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks
Baidu deepbench
Narang, S. and Diamos, G · 2017
Later among the works it cites.
Exploring sparsity in recurrent neural networks
Narang, S., Elsen, E., Diamos, G., and Sengupta, S · 2017
Later among the works it cites.
Plasticine: A reconfigurable architecture for parallel paterns
Prabhakar, R., Zhang, Y., Koeplinger, D., Feldman, M., Zhao, T., Hadjis, S., Pedram, A., Kozyrakis, C., and Olukotun, K · 2017
Later among the works it cites.
A configurable cloud-scale dnn processor for real-time ai
Fowers, J., Ovtcharov, K., Papamichael, M., Massengill, T., Liu, M., Lo, D., Alkalay, S., Haselman, M., Adams, L., Ghandi, M., et al · 2018
Later among the works it cites.
Amc: Automl for model compression and acceleration on mobile devices
He, Y., Lin, J., Liu, Z., Wang, H., Li, L.-J., and Han, S · 2018
Later among the works it cites.
An 787: Intel stratix 10 thermal modeling and management
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, Y.-H., Krishna, T., Emer, J. S., and Sze, V · 2017
Cited alongside, same era.
The intel skylake-x review: Core i9 7900x, i7 7820x and i7 7800x tested
Cutress, I · 2017
Cited alongside, same era.
Inside volta: The world’s most advanced data center gpu
Durant, L., Giroux, O., Harris, M., and Stam, N · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al · 2017
Cited alongside, same era.
https://developer.nvidia.com/cublas
Dense linear algebra on gpus
Cited in the paper.
Dsd: Dense-sparse-dense training for deep neural networks
Han, S., Pool, J., Narang, S., Mao, H., Gong, E., Tang, S., Elsen, E., Vajda, P., Paluri, M., Tran, J., et al
Cited in the paper.
Product specification
Intel
Cited in the paper.
Intel · 2018
Later among the works it cites.
Spatial: A language and compiler for application accelerators
Koeplinger, D., Feldman, M., Prabhakar, R., Zhang, Y., Hadjis, S., Fiszel, R., Zhao, T., Nardi, L., Pedram, A., Kozyrakis, C., and Olukotun, K · 2018
Later among the works it cites.
Nvidia tensor core programmability, performance & precision
Markidis, S., Der Chien, S. W., Laure, E., Peng, I. B., and Vetter, J. S · 2018
Later among the works it cites.
C-lstm: Enabling efficient lstm using structured compression techniques on fpgas
Wang, S., Li, Z., Ding, C., Yuan, B., Qiu, Q., Wang, Y., and Liang, Y · 2018
Later among the works it cites.