Performance estimation of multistreamed, superscalar processors
W. Yamamoto, M. J. Serrano, A. R. Talcott, R. C. Wood, and M. Nemirosky · 1994
Earlier work this paper cites.
Simultaneous multithreading: Maximizing on-chip parallelism
D. M. Tullsen, S. J. Eggers, and H. M. Levy · 1995
Earlier work this paper cites.
Increasing superscalar performance through multistreaming
W. Yamamoto and M. Nemirovsky · 1995
Earlier work this paper cites.
Simultaneous multithreading: A platform for next-generation processors
S. J. Eggers, J. S. Emer, H. M. Levy, J. L. Lo, R. L. Stamm, and D. M. Tullsen · 1997
Earlier work this paper cites.
cuDNN: Efficient primitives for deep learning
Original
S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
S. Han, J. Pool, J. Tran, and W. Dally · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.