Fetching the paper…
Reading the bibliography…
The NVIDIA Volta GPU microarchitecture introduces a specialized unit, called "Tensor Core" that performs one matrix-multiply-and-accumulate on 4x4 matrices per clock cycle.
N. J. Higham, “The accuracy of floating point summation,” SIAM Journal on Scientific Computing , vol. 14, no. 4, pp. 783–799, 1993
1993
Earlier work this paper cites.
A. Buttari, J. Dongarra, J. Langou, J. Langou, P. Luszczek, and J. Kurzak, “Mixed precision iterative refinement techniques for the solution of dense linear systems,” The International Journal of High Performance Computing Applications , vol. 21, no. 4, pp. 457–466, 2007
2007
Earlier work this paper cites.
M. M. Khan, D. R. Lester, L. A. Plana, A. Rast, X. Jin, E. Painkras, and S. B. Furber, “SpiNNaker: mapping neural networks onto a massively-parallel chip multiprocessor,” in Neural Networks, 2008. , 2008, pp. 2849–2856
2008
Earlier work this paper cites.
NVIDIA, “cuBLAS library,” NVIDIA Corporation, Santa Clara, California , vol. 15, no. 27, p. 31, 2008
2008
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
A. Putnam, A. M. Caulfield, E. S. Chung, D. Chiou, K. Constantinides, J. Demme, H. Esmaeilzadeh, J. Fowers, G. P. Gopal, J. Gray et al. , “A reconfigurable fabric for accelerating large-scale datacenter services,” in Computer Architecture (ISCA) . IEEE, 2014, pp. 13–24
2014
Earlier work this paper cites.
P. A. Merolla, J. V. Arthur, R. Alvarez-Icaza, A. S. Cassidy, J. Sawada, F. Akopyan, B. L. Jackson, N. Imam, C. Guo, Y. Nakamura et al. , “A million spiking-neuron integrated circuit with a scalable communication network and interface,” Science , vol. 345, no. 6197, pp. 668–673, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
B. Barry, C. Brick, F. Connor, D. Donohoe, D. Moloney, R. Richmond, M. O’Riordan, and V. Toma, “Always-on vision processing unit for mobile applications,” IEEE Micro , vol. 35, no. 2, pp. 56–66, 2015
2015
Earlier work this paper cites.
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan, “Deep learning with limited numerical precision,” in International Conference on Machine Learning , 2015, pp. 1737–1746
2015
Cited alongside, same era.
S. Markidis, J. Gong, M. Schliephake, E. Laure, A. Hart, D. Henty, K. Heisey, and P. Fischer, “OpenACC acceleration of the Nek5000 spectral element code,” The International Journal of High Performance Computing Applications , vol. 29, no. 3, pp. 311–319, 2015
2015
Cited alongside, same era.
I. Goodfellow, Y. Bengio, and A. Courville, Deep learning . MIT press, 2016
2016
Cited alongside, same era.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard et al. , “TensorFlow: A system for large-scale machine learning.” in OSDI , vol. 16, 2016, pp. 265–283
2016
Cited alongside, same era.
U. Köster, T. Webb, X. Wang, M. Nassar, A. K. Bansal, W. Constable, O. Elibol, S. Hall, L. Hornof, A. Khosrowshahi et al. , “Flexpoint: An adaptive numerical format for efficient training of deep neural networks,” in Advances in Neural Information Processing Systems , 2017, pp. 1740–1750
2017
Later among the works it cites.
2017
Later among the works it cites.
A. Haidar, P. Wu, S. Tomov, and J. Dongarra, “Investigating half precision arithmetic to accelerate dense linear system solvers,” in Proceedings of the 8th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems . ACM, 2017, p. 10
2017
Later among the works it cites.
L. Durant, O. Giroux, M. Harris, and N. Stam, “Inside Volta: The world’s most advanced data center GPU,” 2017, accessed: 2017-01-27. [Online]. Available: https://devblogs.nvidia.com/inside-volta/
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
N. Offermans, O. Marin, M. Schanen, J. Gong, P. Fischer, P. Schlatter, A. Obabko, A. Peplinski, M. Hutchinson, and E. Merzari, “On the strong scaling of the spectral element solver Nek5000 on petascale systems,” in Proceedings of the Exascale Applications and Software Conference 2016 . ACM, 2016, p. 5
2016
Cited alongside, same era.
A. Heinecke, G. Henry, M. Hutchinson, and H. Pabst, “LIBXSMM: accelerating small matrix multiplications by runtime code generation,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2016, p. 84
2016
Cited alongside, same era.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers et al. , “In-datacenter performance analysis of a tensor processing unit,” in Proceedings of the 44th Annual International Symposium on Computer Architecture . ACM, 2017, pp. 1–12
2017
Cited alongside, same era.
K. Carey, “Intel Nervana TM Neural Network Processor: Architecture Update,” 2017, accessed: 2017-01-27. [Online]. Available: https://ai.intel.com/intel-nervana-neural-network-processor-architecture-update/
2017
Cited alongside, same era.
N. Whitehead and A. Fit-Florea, “Precision & performance: Floating point and IEEE 754 compliance for NVIDIA GPUs,” 2011, accessed: 2017-01-27. [Online]. Available: https://developer.download.nvidia.com/assets/cuda/files/NVIDIA-CUDA-Floating-Point.pdf
2017
Cited alongside, same era.
2017
Later among the works it cites.
A. Kerr, D. Merrill, J. Demouth, and J. Tran, “CUTLASS: Fast linear algebra in CUDA C++,” 2017, accessed: 2017-01-27. [Online]. Available: https://devblogs.nvidia.com/cutlass-linear-algebra-cuda/
2017
Later among the works it cites.
J. Dongarra, S. Hammarling, N. J. Higham, S. D. Relton, P. Valero-Lara, and M. Zounon, “The design and performance of batched BLAS on modern high-performance computing systems,” Procedia Computer Science , vol. 108, pp. 495 – 504, 2017, international Conference on Computational Science, ICCS 2017
2017
Later among the works it cites.
C. Cecka, “Low communication FMM-accelerated FFT on GPUs,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . ACM, 2017, p. 54
2017
Later among the works it cites.
P. Luszczek, J. Kurzak, I. Yamazaki, and J. Dongarra, “Towards numerical benchmark for half-precision floating point arithmetic,” in High Performance Extreme Computing Conference (HPEC), 2017 IEEE . IEEE, 2017, pp. 1–5
2017
Later among the works it cites.
NVIDIA, “NVIDIA Tesla V100 GPU architecture,” 2017, accessed: 2018-01-27. [Online]. Available: http://images.nvidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper.pdf
2018
Closest in time.
“NVIDIA CUDA toolkit release notes,” accessed: 2018-03-01. [Online]. Available: http://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html
2018
Closest in time.