Fetching the paper…
Reading the bibliography…
Block Floating Point (BFP) can efficiently support quantization for Deep Neural Network (DNN) training by providing a wide dynamic range via a shared exponent across a group of values.
J. H. Wilkinson, Rounding Errors in Algebraic Processes . Dover Publications, 1964
1964
Earlier work this paper cites.
H. T. Kung, “Why systolic architectures?” IEEE Computer , vol. 15, pp. 37–46, 1982
1982
Earlier work this paper cites.
M. R. Pillmeier, M. J. Schulte, and E. G. Walters III, “Design alternatives for barrel shifters,” in Advanced Signal Processing Algorithms, Architectures, and Implementations XII , vol. 4791. International Society for Optics and Photonics, 2002, pp. 436–447
2002
Earlier work this paper cites.
M. C. Casey, “Single-event effects in digital cmos circuits operating at ultra-low power,” Ph.D. dissertation, Vanderbilt University, 2009, https://www.researchgate.net/publication/224567215_Single-Event_Effects_on_Ultra-Low_Power_CMOS_Circuits
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
M. Everingham and J. Winn, “The pascal visual object classes challenge 2012 (voc2012) development kit,” Pattern Analysis, Statistical Modelling and Computational Learning, Tech. Rep , vol. 8, p. 5, 2011
2011
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Courbariaux, Y. Bengio, and J.-P. David, “Binaryconnect: Training deep neural networks with binary weights during propagations,” in Advances in neural information processing systems , 2015, pp. 3123–3131
2015
Earlier work this paper cites.
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan, “Deep learning with limited numerical precision,” in International Conference on Machine Learning , 2015, pp. 1737–1746
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
P. Judd, J. Albericio, T. Hetherington, T. M. Aamodt, and A. Moshovos, “Stripes: Bit-serial deep neural network computing,” in Microarchitecture (MICRO), 2016 49th Annual IEEE/ACM International Symposium on . IEEE, 2016, pp. 1–12
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. Coleman, D. Narayanan, D. Kang, T. Zhao, J. Zhang, L. Nardi, P. Bailis, K. Olukotun, C. Ré, and M. Zaharia, “Dawnbench: An end-to-end deep learning benchmark and competition,” Training , vol. 100, no. 101, p. 102, 2017
2017
Earlier work this paper cites.
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Quantized neural networks: Training neural networks with low precision weights and activations,” The Journal of Machine Learning Research , vol. 18, no. 1, pp. 6869–6898, 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
E. Park, J. Ahn, and S. Yoo, “Weighted-entropy-based quantization for deep neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 5456–5464
2017
Cited alongside, same era.
X. Sun, J. Choi, C.-Y. Chen, N. Wang, S. Venkataramani, V. V. Srinivasan, X. Cui, W. Zhang, and K. Gopalakrishnan, “Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks,” Advances in neural information processing systems , vol. 32, pp. 4900–4909, 2019
2019
Later among the works it cites.
J. Zhang, X. Chen, M. Song, and T. Li, “Eager pruning: algorithm and architecture support for fast training of deep neural networks,” in 2019 ACM/IEEE 46th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2019, pp. 292–303
2019
Later among the works it cites.
B. Chmiel, L. Ben-Uri, M. Shkolnik, E. Hoffer, R. Banner, and D. Soudry, “Neural gradients are near-lognormal: improved quantized and sparse training,” 2020
2020
Later among the works it cites.
S. Choi, J. Sim, M. Kang, Y. Choi, H. Kim, and L.-S. Kim, “An energy-efficient deep convolutional neural network training accelerator for in situ personalization on smart devices,” IEEE Journal of Solid-State Circuits , vol. 55, no. 10, pp. 2691–2702, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Redmon and A. Farhadi, “Yolo9000: better, faster, stronger,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 7263–7271
2017
Cited alongside, same era.
R. Banner, I. Hubara, E. Hoffer, and D. Soudry, “Scalable methods for 8-bit training of neural networks,” in NeurIPS , 2018, pp. 5151–5159
2018
Cited alongside, same era.
M. Drumond, T. Lin, M. Jaggi, and B. Falsafi, “Training dnns with hybrid block floating point,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems , ser. NIPS’18. Red Hook, NY, USA: Curran Associates Inc., 2018, p. 451–461
2018
Cited alongside, same era.
B. Fleischer, S. Shukla, M. Ziegler, J. Silberman, J. Oh, V. Srinivasan, J. Choi, S. Mueller, A. Agrawal, T. Babinsky, N. Cao, C.-Y. Chen, P. Chuang, T. Fox, G. Gristede, M. Guillorn, H. Haynie, M. Klaiber, D. Lee, S.-H. Lo, G. Maier, M. Scheuermann, S. Venkataramani, C. Vezyrtzis, N. Wang, F. Yee, C. Zhou, P.-F. Lu, B. Curran, L. Chang, and K. Gopalakrishnan, “A scalable multi- teraops deep learning processor core for ai trainina and inference,” in 2018 IEEE Symposium on VLSI Circuits , 2018, pp. 35–36
2018
Cited alongside, same era.
J. Fowers, K. Ovtcharov, M. Papamichael, T. Massengill, M. Liu, D. Lo, S. Alkalay, M. Haselman, L. Adams, M. Ghandi, S. Heil, P. Patel, A. Sapek, G. Weisz, L. Woods, S. Lanka, S. K. Reinhardt, A. M. Caulfield, E. S. Chung, and D. Burger, “A configurable cloud-scale dnn processor for real-time ai,” in Proceedings of the 45th Annual International Symposium on Computer Architecture , ser. ISCA ’18. IEEE Press, 2018, p. 1–14. [Online]. Available: https://doi.org/10.1109/ISCA.2018.00012
2018
Cited alongside, same era.
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 2704–2713
2018
Cited alongside, same era.
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu, “Mixed precision training,” in International Conference on Learning Representations , 2018. [Online]. Available: https://openreview.net/forum?id=r1gs9JgRZ
2018
Cited alongside, same era.
O. Bilaniuk, S. Wagner, Y. Savaria, and J.-P. David, “Bit-slicing fpga accelerator for quantized neural networks,” in 2019 IEEE International Symposium on Circuits and Systems (ISCAS) . IEEE, 2019, pp. 1–5
2019
Cited alongside, same era.
2020
Later among the works it cites.
B. Darvish Rouhani, D. Lo, R. Zhao, M. Liu, J. Fowers, K. Ovtcharov, A. Vinogradsky, S. Massengill, L. Yang, R. Bittner, A. Forin, H. Zhu, T. Na, P. Patel, S. Che, L. Chand Koppaka, X. SONG, S. Som, K. Das, S. T, S. Reinhardt, S. Lanka, E. Chung, and D. Burger, “Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 10 271–10 281. [Online]. Available: https://proceedings.neurips.cc/paper/2020/file/747e32ab0fea7fbd2ad9ec03daa3f840-Paper.pdf
2020
Later among the works it cites.
N. P. Jouppi, D. H. Yoon, G. Kurian, S. Li, N. Patil, J. Laudon, C. Young, and D. Patterson, “A domain-specific supercomputer for training deep neural networks,” Communications of the ACM , vol. 63, no. 7, pp. 67–78, 2020
2020
Later among the works it cites.
M. Mahmoud, I. Edo, A. H. Zadeh, O. M. Awad, G. Pekhimenko, J. Albericio, and A. Moshovos, “Tensordash: Exploiting sparsity to accelerate deep neural network training,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2020, pp. 781–795
2020
Later among the works it cites.
J. Oh, S. Lee, M. Kang, M. Ziegler, J. Silberman, A. Agrawal, S. Venkataramani, B. Fleischer, M. Guillorn, J. Choi, W. Wang, S. Mueller, S. Ben-Yehuda, J. Bonanno, N. Cao, R. Casatuta, C. Chen, M. Cohen, O. Erez, T. Fox, G. Gristede, H. Haynie, V. Ivanov, S. Koswatta, S. Lo, M. Lutz, G. Maier, A. Mesh, Y. Nustov, S. Rider, M. Schaal, M. Scheuermann, X. Sun, N. Wang, F. Yee, C. Zhou, V. Shah, B. Curran, V. Srinivasan, P. Lu, S. Shukla, K. Gopalakrishnan, and L. Chang, “A 3.0 tflops 0.62v scalable processor core for high compute utilization ai training and inference,” in 2020 IEEE Symposium on VLSI Circuits, VLSI Circuits 2020 - Proceedings , ser. IEEE Symposium on VLSI Circuits, Digest of Technical Papers. Institute of Electrical and Electronics Engineers Inc., Jun. 2020, publisher Copyright: © 2020 IEEE.; null ; Conference date: 16-06-2020 Through 19-06-2020
2020
Later among the works it cites.
E. Qin, A. Samajdar, H. Kwon, V. Nadella, S. Srinivasan, D. Das, B. Kaul, and T. Krishna, “Sigma: A sparse and irregular gemm accelerator with flexible interconnects for dnn training,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2020, pp. 58–70
2020
Later among the works it cites.
D. Yang, A. Ghasemazar, X. Ren, M. Golub, G. Lemieux, and M. Lis, “Procrustes: a dataflow and accelerator for sparse deep neural network training,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2020, pp. 711–724
2020
Later among the works it cites.
“Tesla dojo technology,” https://tesla-cdn.thron.com/static/SBY4B9_tesla-dojo-technology_OPNZ0M.pdf?xseo=&response-content-disposition=inline%3Bfilename%3D%22tesla-dojo-technology.pdf%22
2021
Closest in time.
“Bfloat16: The secret to high performance on cloud tpus,” https://cloud.google.com/blog/products/ai-machine-learning/bfloat16-the-secret-to-high-performance-on-cloud-tpus
2021
Closest in time.
“Accelerating ai training with nvidia tf32 tensor cores,” https://developer.nvidia.com/blog/accelerating-ai-training-with-tf32-tensor-cores/
2021
Closest in time.