Fetching the paper…
Reading the bibliography…
During the last two years, the goal of many researchers has been to squeeze the last bit of performance out of HPC system for AI tasks.
1901
Earlier work this paper cites.
1906
Earlier work this paper cites.
A. Petitet, R. Whaley, J. Dongarra, and A. Cleary, “Hpl – a portable implementation of the high-performance linpack benchmark for distributed-memory computers,” - , 01 2008
2008
Earlier work this paper cites.
K. Goto and R. A. Geijn, “Anatomy of high-performance matrix multiplication,” ACM Transactions on Mathematical Software (TOMS) , vol. 34, no. 3, p. 12, 2008
2008
Earlier work this paper cites.
T. Hoefler and A. Lumsdaine, “Message progression in parallel computing - to thread or not to thread?” in 2008 IEEE International Conference on Cluster Computing , 2008, pp. 213–222
2008
Earlier work this paper cites.
V. Aggarwal, Y. Sabharwal, R. Garg, and P. Heidelberger, “Hpcc randomaccess benchmark for next generation supercomputers,” in 2009 IEEE International Symposium on Parallel Distributed Processing , May 2009, pp. 1–11
2009
Earlier work this paper cites.
“Download terabyte click logs,” 12/12/2013. [Online]. Available: https://labs.criteo.com/2013/12/download-terabyte-click-logs
2013
Earlier work this paper cites.
K. Vaidyanathan, K. Pamnany, D. D. Kalamkar, A. Heinecke, M. Smelyanskiy, J. Park, D. Kim, A. Shet, G, B. Kaul, B. Joo, and P. Dubey, “Improving communication performance and scalability of native applications on intel xeon phi coprocessor clusters,” in Proceedings of the 2014 IEEE 28th International Parallel and Distributed Processing Symposium , ser. IPDPS ’14. USA: IEEE Computer Society, 2014, p. 1083–1092. [Online]. Available: https://doi.org/10.1109/IPDPS.2014.113
2014
Earlier work this paper cites.
J. Dongarra, M. A. Heroux, and P. Luszczek, “A new metric for ranking high-performance computing systems,” National Science Review , vol. 3, no. 1, pp. 30–35, 01 2016. [Online]. Available: https://doi.org/10.1093/nsr/nwv084
2016
Earlier work this paper cites.
A. Breuer, A. Heinecke, and M. Bader, “Petascale local time stepping for the ader-dg finite element method,” in 2016 IEEE International Parallel and Distributed Processing Symposium (IPDPS) , 2016, pp. 854–863
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Zhang, J. Yang, and H. Yuen, “Training with low-precision embedding tables,” - , 2018
2018
Cited alongside, same era.
E. Georganas, S. Avancha, K. Banerjee, D. Kalamkar, G. Henry, H. Pabst, and A. Heinecke, “Anatomy of high-performance deep learning convolutions on simd architectures,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis , ser. SC ’18. Piscataway, NJ, USA: IEEE Press, 2018, pp. 66:1–66:12. [Online]. Available: https://doi.org/10.1109/SC.2018.00069
2018
Cited alongside, same era.
2018
Cited alongside, same era.
K. Hazelwood, S. Bird, D. Brooks, S. Chintala, U. Diril, D. Dzhulgakov, M. Fawzy, B. Jia, Y. Jia, A. Kalro, J. Law, K. Lee, J. Lu, P. Noordhuis, M. Smelyanskiy, L. Xiong, and X. Wang, “Applied machine learning at facebook: A datacenter infrastructure perspective,” in 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2018, pp. 620–629
2019
Later among the works it cites.
Facebook, “Dlrm,” https://github.com/facebookresearch/dlrm
2019
Later among the works it cites.
Intel, “oneccl,” https://github.com/oneapi-src/oneCCL
2019
Later among the works it cites.
“Facebook-optimized mlp on multisocket systems,” Accessed on 12/15/2019. [Online]. Available: https://github.com/jspark1105/tbb_test
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
“Tensorflow development summit,” March 30 2018
2018
Cited alongside, same era.
BFLOAT16 – Hardware Numerics Definition . Santa Clara, USA: Intel Corporation, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
N. Dryden, N. Maruyama, T. Moon, T. Benson, M. Snir, and B. Van Essen, “Channel and filter parallelism for large-scale cnn training,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , ser. SC ’19. New York, NY, USA: ACM, 2019, pp. 10:1–10:20. [Online]. Available: http://doi.acm.org/10.1145/3295500.3356207
2019
Cited alongside, same era.
M. Smelyanskiy, “Zion: Facebook next-generation large memory training platform,” in Proceedings of Hot Chips 31 , ser. Hot Chips 31, 2019
2019
Cited alongside, same era.
D. D. Kalamkar, K. Banerjee, S. Srinivasan, S. Sridharan, E. Georganas, M. E. Smorkalov, C. Xu, and A. Heinecke, “Training google neural machine translation on an intel cpu cluster,” in 2019 IEEE International Conference on Cluster Computing (CLUSTER) , Sep. 2019, pp. 1–10
2019
Cited alongside, same era.
2019
Later among the works it cites.
U. Gupta, X. Q. Wang, M. Naumov, C.-J. Wu, B. Reagen, D. Brooks, B. Cottel, K. Hazelwood, B. Jia, H.-H. S. Lee, A. Malevich, D. Mudigere, M. Smelyanskiy, L. Xiong, and X. Zhang, “The architectural implications of facebook’s dnn-based personalized recommendation,” 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) , pp. 488–501, 2019
2019
Later among the works it cites.
NVIDIA, “Hugectr,” https://github.com/NVIDIA/HugeCTR
2019
Later among the works it cites.
E. Georganas, K. Banerjee, D. Kalamkar, S. Avancha, A. Venkat, M. Anderson, G. Henry, H. Pabst, and A. Heinecke, “Harnessing deep learning via a single building block,” 2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS) , 2020
2020
Closest in time.
——, “torch-ccl,” https://github.com/intel/torch-ccl
2020
Closest in time.
Intel Architecture Instruction Set Extensions and Future Features Programming Reference . Santa Clara, USA: Intel Corporation, 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
R. Hwang, T. Kim, Y. Kwon, and M. Rhu, “Centaur: A chiplet-based, hybrid sparse-dense accelerator for personalized recommendations,” ISCA’20 , 2020
2020
Closest in time.