Fetching the paper…
Reading the bibliography…
Training deep neural networks (DNNs) is a computationally expensive job, which can take weeks or months even with high performance GPUs.
1910
Earlier work this paper cites.
K. Kalliojarvi and J. Astola, “Roundoff errors in block-floating-point systems,” IEEE Transactions on Signal Processing (TSP) , vol. 44, no. 4, pp. 783–790, 1996
1996
Earlier work this paper cites.
S. Wilton and N. Jouppi, “CACTI: an enhanced cache access and cycle time model,” IEEE Journal of Solid-State Circuits (JSSC) , vol. 31, no. 5, pp. 677–688, 1996
1996
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2009, pp. 248–255
2009
Earlier work this paper cites.
S. Li, K. Chen, J. H. Ahn, J. B. Brockman, and N. P. Jouppi, “CACTI-P: Architecture-level modeling for SRAM-based structures with advanced leakage reduction techniques,” in IEEE/ACM International Conference on Computer-Aided Design (ICCAD) , 2011, pp. 694–701
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in Neural Information Processing Systems (NIPS) , vol. 25, pp. 1097–1105, 2012
2012
Earlier work this paper cites.
O. Bojar, C. Buck, C. Federmann, B. Haddow, P. Koehn, J. Leveling, C. Monz, P. Pecina, M. Post, H. Saint-Amand, R. Soricut, L. Specia, and A. Tamchyna, “Findings of the 2014 workshop on statistical machine translation,” in Proceedings of the Ninth Workshop on Statistical Machine Translation (WMT) , June 2014, pp. 12–58
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan, “Deep learning with limited numerical precision,” in Proceedings of International Conference on Machine Learning (ICML) , vol. 37, July 2015, pp. 1737–1746
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of International Conference on International Conference on Machine Learning (ICML) , 2015, pp. 448–456
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 770–778
2016
Earlier work this paper cites.
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Binarized neural networks,” in Proceedings of International Conference on Neural Information Processing Systems (NIPS) , 2016, pp. 4114–4122
2016
Earlier work this paper cites.
J. Kung, D. Kim, and S. Mukhopadhyay, “Dynamic approximation with feedback control for energy-efficient recurrent neural network hardware,” in Proceedings of the International Symposium on Low Power Electronics and Design (ISLPED) , 2016, pp. 168–173
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. D. Lin, S. S. Talathi, and V. S. Annapureddy, “Fixed point quantization of deep convolutional networks,” in Proceedings of International Conference on Machine Learning (ICML) , 2016, pp. 2849–2858
2016
Earlier work this paper cites.
B. Moons, B. D. Brabandere, L. V. Gool, and M. Verhelst, “Energy-efficient convnets through approximate computing,” in IEEE Winter Conference on Applications of Computer Vision (WACV) , Mar. 2016
2016
Earlier work this paper cites.
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “XNOR-Net: ImageNet classification using binary convolutional neural networks,” in European Conference on Computer Vision (ECCV) . Springer International Publishing, 2016, pp. 525–542
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Synopsys, “IC Compiler II Implementation User Guide: Version L-2016.03,” 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon, “In-datacenter performance analysis of a tensor processing unit,” SIGARCH Comput. Archit. News , vol. 45, no. 2, pp. 1–12, 2017
2017
Earlier work this paper cites.
U. Köster, T. J. Webb, X. Wang, M. Nassar, A. K. Bansal, W. H. Constable, O. H. Elibol, S. Gray, S. Hall, L. Hornof, A. Khosrowshahi, C. Kloss, R. J. Pai, and N. Rao, “Flexpoint: An adaptive numerical format for efficient training of deep neural networks,” in Proceedings of International Conference on Neural Information Processing Systems (NIPS) , 2017, pp. 1740–1750
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
B. Moons, R. Uytterhoeven, W. Dehaene, and M. Verhelst, “Envision: A 0.26-to-10tops/w subword-parallel dynamic-voltage-accuracy-frequency-scalable convolutional neural network processor in 28nm FDSOI,” in IEEE International Solid-State Circuits Conference (ISSCC) , 2017, pp. 246–247
2017
Cited alongside, same era.
D. Shin, J. Lee, J. Lee, and H.-J. Yoo, “DNPU: An 8.1TOPS/W reconfigurable CNN-RNN processor for general-purpose deep neural networks,” in IEEE International Solid-State Circuits Conference (ISSCC) , 2017, pp. 240–241
G. Huang, Z. Liu, G. Pleiss, L. Van Der Maaten, and K. Weinberger, “Convolutional networks with dense connectivity,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , pp. 1–1, 2019
2019
Later among the works it cites.
W. Jung, D. Jung, B. Kim, S. Lee, W. Rhee, and J. H. Ahn, “Restructuring batch normalization to accelerate CNN training,” in Proceedings of the Conference on Systems and Machine Learning (SysML) , 2019, pp. 1–13
2019
Later among the works it cites.
J. Lee, J. Lee, D. Han, J. Lee, G. Park, and H.-J. Yoo, “LNPU: A 25.3TFLOPS/W sparse deep-neural-network learning processor with fine-grained mixed precision of FP8-FP16,” in IEEE International Solid- State Circuits Conference (ISSCC) , 2019, pp. 142–144
2019
Later among the works it cites.
S. Ryu, H. Kim, W. Yi, and J.-J. Kim, “BitBlade: Area and energy-efficient precision-scalable neural network accelerator with bitwise summation,” in ACM/IEEE Design Automation Conference (DAC) , 2019, pp. 1–6
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Synopsys, “Design Compiler User Guide: Version N-2017.09,” 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NeurIPS) , vol. 30, 2017
2017
Cited alongside, same era.
S. Venkataramani, A. Ranjan, S. Banerjee, D. Das, S. Avancha, A. Jagannathan, A. Durg, D. Nagaraj, B. Kaul, P. Dubey, and A. Raghunathan, “Scaledeep: A scalable compute architecture for learning and evaluating deep networks,” in ACM/IEEE International Symposium on Computer Architecture (ISCA) , 2017, pp. 13–26
2017
Cited alongside, same era.
R. Banner, I. Hubara, E. Hoffer, and D. Soudry, “Scalable methods for 8-bit training of neural networks,” in Proceedings of International Conference on Neural Information Processing Systems (NeurIPS) , 2018, pp. 5151–5159
2018
Cited alongside, same era.
2018
Cited alongside, same era.
M. Drumond, T. Lin, M. Jaggi, and B. Falsafi, “Training DNNs with hybrid block floating point,” in Advances in Neural Information Processing Systems (NeurIPS) , 2018
2018
Cited alongside, same era.
T. Geng, T. Wang, A. Sanaullah, C. Yang, R. Xu, R. Patel, and M. Herbordt, “FPDeep: Acceleration and load balancing of CNN training on FPGA clusters,” in IEEE International Symposium on Field-Programmable Custom Computing Machines (FCCM) , 2018, pp. 81–84
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Later among the works it cites.
X. Sun, J. Choi, C.-Y. Chen, N. Wang, S. Venkataramani, V. V. Srinivasan, X. Cui, W. Zhang, and K. Gopalakrishnan, “Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks,” in Advances in Neural Information Processing Systems (NeurIPS) , H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, Eds., vol. 32, 2019
2019
Later among the works it cites.
M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proceedings of the International Conference on Machine Learning (ICML) , vol. 97, June 2019, pp. 6105–6114
2019
Later among the works it cites.
M. Andrychowicz, B. Baker, M. Chociej, R. Józefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba, “Learning dexterous in-hand manipulation,” The International Journal of Robotics Research (IJRR) , vol. 39, no. 1, pp. 3–20, 2020
2020
Later among the works it cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in Advances in Neural Information Processing Systems (NeurIPS) , 2020
2020
Later among the works it cites.
B. Hickmann, J. Chen, M. Rotzin, A. Yang, M. Urbanski, and S. Avancha, “Intel Nervana neural network processor-t (NNP-T) fused floating point many-term dot product,” in IEEE Symposium on Computer Arithmetic (ARITH) , 2020, pp. 133–136
2020
Later among the works it cites.
J. Lin, W.-M. Chen, Y. Lin, J. Cohn, C. Gan, and S. Han, “MCUNet: Tiny deep learning on IoT devices,” in Advances in Neural Information Processing Systems (NeurIPS) , vol. 33. Curran Associates, Inc., 2020, pp. 11 711–11 722
2020
Later among the works it cites.
E. Qin, A. Samajdar, H. Kwon, V. Nadella, S. Srinivasan, D. Das, B. Kaul, and T. Krishna, “SIGMA: A sparse and irregular GEMM accelerator with flexible interconnects for DNN training,” in IEEE International Symposium on High Performance Computer Architecture (HPCA) , Feb. 2020, pp. 58–70
2020
Later among the works it cites.
X. Sun, N. Wang, C.-Y. Chen, J. Ni, A. Agrawal, X. Cui, S. Venkataramani, K. El Maghraoui, V. V. Srinivasan, and K. Gopalakrishnan, “Ultra-low precision 4-bit training of deep neural networks,” in Advances in Neural Information Processing Systems (NeurIPS) , vol. 33, 2020, pp. 1796–1807
2020
Later among the works it cites.
J. Verbraeken, M. Wolting, J. Katzy, J. Kloppenburg, T. Verbelen, and J. S. Rellermeyer, “A survey on distributed machine learning,” ACM Comput. Surv. , vol. 53, no. 2, Mar. 2020. [Online]. Available: https://doi.org/10.1145/3377454
2020
Later among the works it cites.
D. Yang, A. Ghasemazar, X. Ren, M. Golub, G. Lemieux, and M. Lis, “Procrustes: a dataflow and accelerator for sparse deep neural network training,” in Proceedings of the IEEE/ACM International Symposium on Microarchitecture (MICRO) , Oct. 2020
2020
Later among the works it cites.
A. Agrawal, S. K. Lee, J. Silberman, M. Ziegler, M. Kang, S. Venkataramani, N. Cao, B. Fleischer, M. Guillorn, M. Cohen, S. Mueller, J. Oh, M. Lutz, J. Jung, S. Koswatta, C. Zhou, V. Zalani, J. Bonanno, R. Casatuta, C.-Y. Chen, J. Choi, H. Haynie, A. Herbert, R. Jain, M. Kar, K.-H. Kim, Y. Li, Z. Ren, S. Rider, M. Schaal, K. Schelm, M. Scheuermann, X. Sun, H. Tran, N. Wang, W. Wang, X. Zhang, V. Shah, B. Curran, V. Srinivasan, P.-F. Lu, S. Shukla, L. Chang, and K. Gopalakrishnan, “9.1 a 7nm 4-core AI chip with 25.6TFLOPS hybrid FP8 training, 102.4TOPS INT4 inference and workload-aware throttling,” in IEEE International Solid- State Circuits Conference (ISSCC) , vol. 64, 2021, pp. 144–146
2021
Later among the works it cites.
Google, “Cloud TPU,” https://cloud.google.com/tpu
2021
Later among the works it cites.
Google Cloud, “BFloat16: The secret to high performance on cloud TPUs,” https://cloud.google.com/blog/products/ai-machine-learning/bfloat16-the-secret-to-high-performance-on-cloud-tpus
2021
Later among the works it cites.
Y. Jang, S. Kim, D. Kim, S. Lee, and J. Kung, “Deep partitioned training from near-storage computing to DNN accelerators,” IEEE Computer Architecture Letters (CAL) , pp. 1–4, 2021
2021
Later among the works it cites.
A. Krizhevsky, V. Nair, and G. Hinton, “CIFAR-10 and CIFAR-100 dataset,” https://www.cs.toronto.edu/~kriz/cifar.html
2021
Later among the works it cites.
A. Mirhoseini, A. Goldie, M. Yazgan, J. W. Jiang, E. Songhori, S. Wang, Y.-J. Lee, E. Johnson, O. Pathak, A. Nazi, J. Pak, A. Tong, K. Srinivasa, W. Hang, E. Tuncer, Q. V. Le, J. Laudon, R. Ho, R. Carpenter, and J. Dean, “A graph placement methodology for fast chip design,” Nature , vol. 594, pp. 207–212, June 2021. [Online]. Available: https://doi.org/10.1038/s41586-021-03544-w
2021
Later among the works it cites.
J. Park, S. Lee, and D. Jeon, “A 40nm 4.81TFLOPS/W 8b floating-point training processor for non-sparse neural networks using shared exponent bias and 24-way fused multiply-add tree,” in IEEE International Solid- State Circuits Conference (ISSCC) , vol. 64, 2021, pp. 1–3
2021
Later among the works it cites.
S. Venkataramani, V. Srinivasan, W. Wang, S. Sen, J. Zhang, A. Agrawal, M. Kar, S. Jain, A. Mannari, H. Tran, Y. Li, E. Ogawa, K. Ishizaki, H. Inoue, M. Schaal, M. Serrano, J. Choi, X. Sun, N. Wang, C.-Y. Chen, A. Allain, J. Bonano, N. Cao, R. Casatuta, M. Cohen, B. Fleischer, M. Guillorn, H. Haynie, J. Jung, M. Kang, K.-h. Kim, S. Koswatta, S. Lee, M. Lutz, S. Mueller, J. Oh, A. Ranjan, Z. Ren, S. Rider, K. Schelm, M. Scheuermann, J. Silberman, J. Yang, V. Zalani, X. Zhang, C. Zhou, M. Ziegler, V. Shah, M. Ohara, P.-F. Lu, B. Curran, S. Shukla, L. Chang, and K. Gopalakrishnan, “RaPiD: AI accelerator for ultra-low precision training and inference,” in ACM/IEEE International Symposium on Computer Architecture (ISCA) , 2021, pp. 153–166
2021
Later among the works it cites.
J.-H. Yoon, M. Chang, W.-S. Khwa, Y.-D. Chih, M.-F. Chang, and A. Raychowdhury, “A 40nm 64Kb 56.67TOPS/W read-disturb-tolerant compute-in-memory/digital RRAM macro with active-feedback-based read and in-situ write verification,” in IEEE International Solid- State Circuits Conference (ISSCC) , vol. 64, 2021, pp. 404–406
2021
Later among the works it cites.