Fetching the paper…
Reading the bibliography…
In this paper, we provide a deep dive into the deployment of inference accelerators at Facebook.
1906
Earlier work this paper cites.
1911
Earlier work this paper cites.
S. Rendle, “Factorization machines,” in 2010 IEEE International Conference on Data Mining , 2010, pp. 995–1000
2010
Earlier work this paper cites.
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun et al. , “Dadiannao: A machine-learning supercomputer,” in 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture . IEEE, 2014, pp. 609–622
2014
Earlier work this paper cites.
X. He, J. Pan, O. Jin, T. Xu, B. Liu, T. Xu, Y. Shi, A. Atallah, R. Herbrich, S. Bowers, and J. Q. n. Candela, “Practical lessons from predicting clicks on ads at facebook,” in Proceedings of the Eighth International Workshop on Data Mining for Online Advertising , ser. ADKDD’14. New York, NY, USA: Association for Computing Machinery, 2014, p. 1–9. [Online]. Available: https://doi.org/10.1145/2648584.2648589
2014
Earlier work this paper cites.
Chen, Yu-Hsin and Krishna, Tushar and Emer, Joel and Sze, Vivienne, “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,” in IEEE International Solid-State Circuits Conference, ISSCC 2016, Digest of Technical Papers , 2016, pp. 262–263
2016
Earlier work this paper cites.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “Eie: Efficient inference engine on compressed deep neural network,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) , 2016, pp. 243–254
2016
Earlier work this paper cites.
S. Liu, Z. Du, J. Tao, D. Han, T. Luo, Y. Xie, Y. Chen, and T. Chen, “Cambricon: An instruction set architecture for neural networks,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2016, pp. 393–405
2016
Earlier work this paper cites.
B. Reagen, P. Whatmough, R. Adolf, S. Rama, H. Lee, S. K. Lee, J. M. Hernández-Lobato, G. Wei, and D. Brooks, “Minerva: Enabling low-power, highly-accurate deep neural network accelerators,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) , 2016, pp. 267–278
2016
Earlier work this paper cites.
H. Sharma, J. Park, D. Mahajan, E. Amaro, J. K. Kim, C. Shao, A. Mishra, and H. Esmaeilzadeh, “From high-level deep neural models to fpgas,” in 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2016, pp. 1–12
2016
Earlier work this paper cites.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 5987–5995
2017
Earlier work this paper cites.
F. Borisyuk, A. Gordo, and V. Sivakumar, “Rosetta: Large scale system for text detection and recognition in images,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , ser. KDD ’18. New York, NY, USA: Association for Computing Machinery, 2018, p. 71–79. [Online]. Available: https://doi.org/10.1145/3219819.3219861
2018
Earlier work this paper cites.
E. Chung, J. Fowers, K. Ovtcharov, M. Papamichael, A. Caulfield, T. Massengill, M. Liu, M. Ghandi, D. Lo, S. Reinhardt, S. Alkalay, H. Angepat, D. Chiou, A. Forin, D. Burger, L. Woods, G. Weisz, M. Haselman, and D. Zhang, “Serving dnns in real time at datacenter scale with project brainwave,” IEEE Micro , vol. 38, pp. 8–20, March 2018. [Online]. Available: https://www.microsoft.com/en-us/research/publication/serving-dnns-real-time-datacenter-scale-project-brainwave/
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Fowers, K. Ovtcharov, M. Papamichael, T. Massengill, M. Liu, D. Lo, S. Alkalay, M. Haselman, L. Adams, M. Ghandi, S. Heil, P. Patel, A. Sapek, G. Weisz, L. Woods, S. Lanka, S. Reinhardt, A. Caulfield, E. Chung, and D. Burger, “A configurable cloud-scale dnn processor for real-time ai,” in Proceedings of the 45th International Symposium on Computer Architecture, 2018 . ACM, June 2018. [Online]. Available: https://www.microsoft.com/en-us/research/publication/a-configurable-cloud-scale-dnn-processor-for-real-time-ai/
2018
Earlier work this paper cites.
K. Hazelwood, S. Bird, D. Brooks, S. Chintala, U. Diril, D. Dzhulgakov, M. Fawzy, B. Jia, Y. Jia, A. Kalro, J. Law, K. Lee, J. Lu, P. Noordhuis, M. Smelyanskiy, L. Xiong, and X. Wang, “Applied machine learning at facebook: A datacenter infrastructure perspective,” in 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2018, pp. 620–629
2018
Earlier work this paper cites.
H. Kwon, A. Samajdar, and T. Krishna, “Maeri: Enabling flexible dataflow mapping over dnn accelerators via reconfigurable interconnects,” ACM SIGPLAN Notices , vol. 53, no. 2, pp. 461–475, 2018
2018
Earlier work this paper cites.
D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. van der Maaten, “Exploring the limits of weakly supervised pretraining,” 2018
2018
Earlier work this paper cites.
J. Park, M. Naumov, P. Basu, S. Deng, A. Kalaiah, D. Khudia, J. Law, P. Malani, A. Malevich, S. Nadathur, J. Pino, M. Schatz, A. Sidorov, V. Sivakumar, A. Tulloch, X. Wang, Y. Wu, H. Yuen, U. Diril, D. Dzhulgakov, K. Hazelwood, B. Jia, Y. Jia, L. Qiao, V. Rao, N. Rotem, S. Yoo, and M. Smelyanskiy, “Deep learning inference in facebook data centers: Characterization, performance optimizations and hardware implications,” 2018
2018
Cited alongside, same era.
A. Adcock, V. Reis, M. Singh, Z. Yan, L. van der Maaten, K. Zhang, S. Motwani, J. Guerin, N. Goyal, I. Misra, L. Gustafson, C. Changhan, and P. Goyal, “Classy vision,” 2019
2019
Cited alongside, same era.
Y. Chen, H. Fan, B. Xu, Z. Yan, Y. Kalantidis, M. Rohrbach, S. Yan, and J. Feng, “Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2019
2019
Cited alongside, same era.
U. Gupta, C. Wu, X. Wang, M. Naumov, B. Reagen, D. Brooks, B. Cottel, K. Hazelwood, M. Hempstead, B. Jia, H. S. Lee, A. Malevich, D. Mudigere, M. Smelyanskiy, L. Xiong, and X. Zhang, “The architectural implications of facebook’s dnn-based personalized recommendation,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2020, pp. 488–501
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Hsia, U. Gupta, M. Wilkening, C.-J. Wu, G.-Y. Wei, and D. Brooks, “Cross-stack workload characterization of deep recommendation systems,” 2020
2020
Later among the works it cites.
R. Hwang, T. Kim, Y. Kwon, and M. Rhu, “Centaur: A chiplet-based, hybrid sparse-dense accelerator for personalized recommendations,” 2020
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Jang, J. Kim, J.-E. Jo, J. Lee, and J. Kim, “Mnnfast: A fast and scalable system architecture for memory-augmented neural networks,” in Proceedings of the 46th International Symposium on Computer Architecture , ser. ISCA ’19. New York, NY, USA: Association for Computing Machinery, 2019, p. 250–263. [Online]. Available: https://doi.org/10.1145/3307650.3322214
2019
Cited alongside, same era.
Z. Jia, O. Padon, J. Thomas, T. Warszawski, M. Zaharia, and A. Aiken, “Taso: optimizing deep learning computation with automatic generation of graph substitutions,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles , 2019, pp. 47–62
2019
Cited alongside, same era.
D. Kalamkar, D. Mudigere, N. Mellempudi, D. Das, K. Banerjee, S. Avancha, D. T. Vooturi, N. Jammalamadaka, J. Huang, H. Yuen, J. Yang, J. Park, A. Heinecke, E. Georganas, S. Srinivasan, A. Kundu, M. Smelyanskiy, B. Kaul, and P. Dubey, “A study of bfloat16 for deep learning training,” 2019
2019
Cited alongside, same era.
V. R. Kevin Lee. Accelerating facebook’s infrastructure with application-specific hardware. [Online]. Available: https://engineering.fb.com/2019/03/14/data-center-engineering/accelerating-infrastructure/
2019
Cited alongside, same era.
Y. Kwon, Y. Lee, and M. Rhu, “Tensordimm: A practical near-memory processing architecture for embeddings and tensor operations in deep learning,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , 2019, pp. 740–753
2019
Cited alongside, same era.
N. Rotem, J. Fix, S. Abdulrasool, G. Catron, S. Deng, R. Dzhabarov, N. Gibson, J. Hegeman, M. Lele, R. Levenstein, J. Montgomery, B. Maher, S. Nadathur, J. Olesen, J. Park, A. Rakhov, M. Smelyanskiy, and M. Wang, “Glow: Graph lowering compiler techniques for neural networks,” 2019
2019
Cited alongside, same era.
J. R. Stevens, A. Ranjan, D. Das, B. Kaul, and A. Raghunathan, “Manna: An accelerator for memory-augmented neural networks,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , ser. MICRO ’52. New York, NY, USA: Association for Computing Machinery, 2019, p. 794–806. [Online]. Available: https://doi.org/10.1145/3352460.3358304
2019
Cited alongside, same era.
D. Tran, H. Wang, L. Torresani, and M. Feiszli, “Video classification with channel-separated convolutional networks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2019
2019
Cited alongside, same era.
E. Baek, D. Kwon, and J. Kim, “A multi-neural network acceleration architecture,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2020, pp. 940–953
2020
Cited alongside, same era.
Later among the works it cites.
N. P. Jouppi, D. H. Yoon, G. Kurian, S. Li, N. Patil, J. Laudon, C. Young, and D. Patterson, “A domain-specific supercomputer for training deep neural networks,” Commun. ACM , vol. 63, no. 7, p. 67–78, Jun. 2020. [Online]. Available: https://doi.org/10.1145/3360307
2020
Later among the works it cites.
L. Ke, U. Gupta, B. Y. Cho, D. Brooks, V. Chandra, U. Diril, A. Firoozshahian, K. Hazelwood, B. Jia, H.-H. S. Lee et al. , “Recnmp: Accelerating personalized recommendation with near-memory processing,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2020, pp. 790–803
2020
Later among the works it cites.
D. Kwon, S. Hur, H. Jang, E. Nurvitadhi, and J. Kim, “Scalable multi-fpga acceleration for large rnns with full parallelism levels,” in 2020 57th ACM/IEEE Design Automation Conference (DAC) , 2020, pp. 1–6
2020
Later among the works it cites.
H. Kwon, L. Lai, M. Pellauer, T. Krishna, Y.-H. Chen, and V. Chandra, “Heterogeneous dataflow accelerators for multi-dnn workloads,” 2020
2020
Later among the works it cites.
M. Lui, Y. Yetim, Özgür Özkan, Z. Zhao, S.-Y. Tsai, C.-J. Wu, and M. Hempstead, “Understanding capacity-driven scale-out neural recommendation inference,” 2020
2020
Later among the works it cites.
P. Mattson, V. J. Reddi, C. Cheng, C. Coleman, G. Diamos, D. Kanter, P. Micikevicius, D. Patterson, G. Schmuelling, H. Tang, G. Wei, and C. Wu, “Mlperf: An industry standard benchmark suite for machine learning performance,” IEEE Micro , vol. 40, no. 2, pp. 8–16, 2020
2020
Later among the works it cites.
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollár, “Designing network design spaces,” 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
L. Song, F. Chen, Y. Zhuo, X. Qian, H. Li, and Y. Chen, “Accpar: Tensor partitioning for heterogeneous deep learning accelerators,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2020, pp. 342–355
2020
Later among the works it cites.
W. Wang, D. Tran, and M. Feiszli, “What makes training multi-modal classification networks hard?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020
2020
Later among the works it cites.
M. Ye, D. Choudhary, J. Yu, E. Wen, Z. Chen, J. Yang, J. Park, Q. Liu, and A. Kejariwal, “Adaptive dense-to-sparse paradigm for pruning online recommendation system with non-stationary data,” 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Park, J. Jang, S. Kim, B. Na, and S. Yoon, “Memory-augmented neural networks on fpga for real-time and energy-efficient question answering,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , vol. 29, no. 1, pp. 162–175, 2021
2021
Closest in time.