Fetching the paper…
Reading the bibliography…
Why systolic architectures?
Kung, H.-T · 1982
Earlier work this paper cites.
A parallel ieee p754 decimal floating-point multiplier
Hickmann, B., Krioukov, A., Schulte, M., and Erle, M · 2007
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Orion 2.0: A fast and accurate noc power and area model for early-stage design space exploration
Kahng, A. B., Li, B., Peh, L.-S., and Samadi, K · 2009
Earlier work this paper cites.
Understanding the energy consumption of dynamic random access memories
Vogelsang, T · 2010
Earlier work this paper cites.
Cacti-3dd: Architecture-level modeling for 3d die-stacked dram main memory
Chen, K., Li, S., Muralimanohar, N., Ahn, J. H., Brockman, J. B., and Jouppi, N. P · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Autotuning gemm kernels for the fermi gpu
Kurzak, J., Tomov, S., and Dongarra, J · 2012
Earlier work this paper cites.
Minimizing energy of integer unit by higher voltage flip-flop: Vddmin-aware dual supply voltage technique
Fuketa, H., Hirairi, K., Yasufuku, T., Takamiya, M., Nomura, M., Shinohara, H., and Sakurai, T · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
Chetlur, S., Woolley, C., Vandermersch, P., Cohen, J., Tran, J., Catanzaro, B., and Shelhamer, E · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., and Darrell, T · 2014
Earlier work this paper cites.
27.8 a static contention-free single-phase-clocked 24t flip-flop in 45nm for low-power applications
Kim, Y., Jung, W., Lee, I., Dong, Q., Henry, M., Sylvester, D., and Blaauw, D · 2014
Cited alongside, same era.
Efficient mini-batch training for stochastic optimization
Li, M., Zhang, T., Chen, Y., and Smola, A. J · 2014
Cited alongside, same era.
Shidiannao: Shifting vision processing closer to the sensor
Du, Z., Fasthuber, R., Chen, T., Ienne, P., Li, L., Luo, T., Feng, X., Chen, Y., and Temam, O · 2015
Cited alongside, same era.
Han, S., Mao, H., and Dally, W. J · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Machine learning for systems and systems for machine learning
Dean, J · 2017
Later among the works it cites.
Tetris: Scalable and efficient neural network acceleration with 3d memory
Gao, M., Pu, J., Yang, X., Horowitz, M., and Kozyrakis, C · 2017
Later among the works it cites.
Accurate, large minibatch sgd: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Later among the works it cites.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al · 2017
Later among the works it cites.
Flexflow: A flexible dataflow accelerator architecture for convolutional neural networks
Lu, W., Yan, G., Li, J., Gong, S., Han, Y., and Li, X · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Cited alongside, same era.
Fused-layer cnn accelerators
Alwani, M., Chen, H., Ferdman, M., and Milder, P · 2016
Cited alongside, same era.
Distributed deep learning using synchronous stochastic gradient descent
Das, D., Avancha, S., Mudigere, D., Vaidynathan, K., Sridharan, S., Kalamkar, D., Kaul, B., and Dubey, P · 2016
Cited alongside, same era.
Eie: efficient inference engine on compressed deep neural network
Han, S., Liu, X., Mao, H., Pu, J., Pedram, A., Horowitz, M. A., and Dally, W. J · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Joint Electron Device Engineering Council, Jan. 2016
High Bandwidth Memory (HBM) DRAM, JESD235A · 2016
Cited alongside, same era.
Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks
Chen, Y.-H., Krishna, T., Emer, J. S., and Sze, V · 2017
Cited alongside, same era.
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaev, O., Venkatesh, G., et al · 2017
Later among the works it cites.
Nvidia tesla v100 gpu architecture
nvidia · 2017
Later among the works it cites.
Scnn: An accelerator for compressed-sparse convolutional neural networks
Parashar, A., Rhu, M., Mukkara, A., Puglielli, A., Venkatesan, R., Khailany, B., Emer, J., Keckler, S. W., and Dally, W. J · 2017
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A. A · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Closest in time.
Wu, Y. and He, K · 2018
Closest in time.