Fetching the paper…
Reading the bibliography…
We devise a performance model for GPU training of Deep Learning Recommendation Models (DLRM), whose GPU utilization is low compared to other well-optimized CV and NLP models.
1906
Earlier work this paper cites.
S. H. Walker and D. B. Duncan, “Estimation of the probability of an event as a function of several independent variables,” Biometrika , vol. 54, no. 1–2, pp. 167–179, Jun. 1967. [Online]. Available: https://doi.org/10.1093/biomet/54.1-2.167
1967
Earlier work this paper cites.
J. L. Herlocker, J. A. Konstan, and J. Riedl, “Explaining collaborative filtering recommendations,” in Proceedings of the 2000 ACM Conference on Computer Supported Cooperative Work , ser. CSCW ’00, 2000, pp. 241–250. [Online]. Available: https://doi.org/10.1145/358916.358995
2000
Earlier work this paper cites.
S. Williams, A. Waterman, and D. Patterson, “Roofline: An insightful visual performance model for multicore architectures,” Communications of the ACM , vol. 52, no. 4, pp. 65–76, 2009
2009
Earlier work this paper cites.
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” in Proceedings of the 22nd ACM International Conference on Multimedia , ser. MM ’14, 2014, pp. 675–678. [Online]. Available: https://doi.org/10.1145/2647868.2654889
2014
Earlier work this paper cites.
X. Ning, C. Desrosiers, and G. Karypis, “A comprehensive survey of neighborhood-based recommendation methods,” in Recommender Systems Handbook , F. Ricci, L. Rokach, and B. Shapira, Eds. Springer, 2015, pp. 37–76. [Online]. Available: https://doi.org/10.1007/978-1-4899-7637-6_2
2015
Earlier work this paper cites.
J. Gómez-Luna, I.-J. Sung, L.-W. Chang, J. M. González-Linares, N. Guil, and W.-M. W. Hwu, “In-place matrix transposition on GPUs,” IEEE Transactions on Parallel and Distributed Systems , vol. 27, no. 3, pp. 776–788, 2016. [Online]. Available: https://doi.org/10.1109/TPDS.2015.2412549
2015
Earlier work this paper cites.
P. Covington, J. Adams, and E. Sargin, “Deep neural networks for YouTube recommendations,” in Proceedings of the 10th ACM Conference on Recommender Systems , ser. RecSys ’16, Sep. 2016, pp. 191–198. [Online]. Available: https://doi.org/10.1145/2959100.2959190
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition , ser. CVPR 2016, Jun. 2016, pp. 770–778. [Online]. Available: https://doi.org/10.1109/CVPR.2016.90
2016
Earlier work this paper cites.
H.-T. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir, R. Anil, Z. Haque, L. Hong, V. Jain, X. Liu, and H. Shah, “Wide & deep learning for recommender systems,” in Proceedings of the 1st Workshop on Deep Learning for Recommender Systems , ser. DLRS 2016, 2016, pp. 7–10. [Online]. Available: https://doi.org/10.1145/2988450.2988454
2016
Earlier work this paper cites.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: A system for large-scale machine learning,” in 12th USENIX Symposium on Operating Systems Design and Implementation , ser. OSDI 16, 2016, pp. 265–283. [Online]. Available: https://www.usenix.org/system/files/conference/osdi16/osdi16-abadi.pdf
2016
Earlier work this paper cites.
B. Smith and G. Linden, “Two decades of recommender systems at Amazon.com,” IEEE Internet Computing , vol. 21, no. 3, pp. 12–18, May/Jun. 2017. [Online]. Available: https://doi.org/10.1109/MIC.2017.72
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds., vol. 30. Curran Associates, Inc., Dec. 2017. [Online]. Available: https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf
2017
Cited alongside, same era.
H. Guo, R. Tang, Y. Ye, Z. Li, and X. He, “DeepFM: A factorization-machine based neural network for CTR prediction,” in Proceedings of the 26th International Joint Conference on Artificial Intelligence , ser. IJCAI’17. AAAI Press, 2017, pp. 1725–1731
2017
Cited alongside, same era.
R. Wang, B. Fu, G. Fu, and M. Wang, “Deep & cross network for ad click predictions,” in Proceedings of the ADKDD’17 , ser. ADKDD’17, 2017. [Online]. Available: https://doi.org/10.1145/3124749.3124754
2017
Cited alongside, same era.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “PyTorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems 33 , ser. NeurIPS 2019, Dec. 2019, pp. 8026–8037. [Online]. Available: https://proceedings.neurips.cc/paper/2019/file/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf
2019
Later among the works it cites.
Facebook. (2019, Jun.) DLRM Github repo. [Online]. Available: https://github.com/facebookresearch/dlrm
2019
Later among the works it cites.
E. Elahi and A. Chandrashekar, “Learning representations of hierarchical slates in collaborative filtering,” in Fourteenth ACM Conference on Recommender Systems , ser. RecSys ’20, Sep. 2020, pp. 703–707. [Online]. Available: https://doi.org/10.1145/3383313.3418484
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Konstantinidis and Y. Cotronis, “A quantitative roofline model for GPU kernel performance estimation using micro-benchmarks and hardware metric profiling,” Journal of Parallel and Distributed Computing , vol. 107, pp. 37–56, 2017. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0743731517301247
2017
Cited alongside, same era.
J. Lian, X. Zhou, F. Zhang, Z. Chen, X. Xie, and G. Sun, “XDeepFM: Combining explicit and implicit feature interactions for recommender systems,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , ser. KDD ’18, 2018, pp. 1754–1763. [Online]. Available: https://doi.org/10.1145/3219819.3220023
2018
Cited alongside, same era.
G. Zhou, X. Zhu, C. Song, Y. Fan, H. Zhu, X. Ma, Y. Yan, J. Jin, H. Li, and K. Gai, “Deep interest network for click-through rate prediction,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , ser. KDD ’18, 2018, pp. 1059–1068. [Online]. Available: https://doi.org/10.1145/3219819.3219823
2018
Cited alongside, same era.
D. Justus, J. Brennan, S. Bonner, and A. S. McGough, “Predicting the computational cost of deep learning models,” in 2018 IEEE International Conference on Big Data , ser. BigData 2018, Dec. 2018, pp. 3873–3882. [Online]. Available: https://doi.org/10.1109/BigData.2018.8622396
2018
Cited alongside, same era.
J. Vedurada, A. Suresh, A. S. Rajam, J. Kim, C. Hong, A. Panyala, S. Krishnamoorthy, V. K. Nandivada, R. K. Srivastava, and P. Sadayappan, “TTLG - an efficient tensor transposition library for GPUs,” in 2018 IEEE International Parallel and Distributed Processing Symposium (IPDPS) , 2018, pp. 578–588. [Online]. Available: https://doi.org/10.1109/IPDPS.2018.00067
2018
Cited alongside, same era.
M. Tennenholtz and O. Kurland, “Rethinking search engines and recommendation systems: A game theoretic perspective,” Commun. ACM , vol. 62, no. 12, pp. 66–75, Nov. 2019. [Online]. Available: https://doi.org/10.1145/3340922
2019
Cited alongside, same era.
W. Zhao, J. Zhang, D. Xie, Y. Qian, R. Jia, and P. Li, “AIBox: CTR prediction model training on a single node,” in Proceedings of the 28th ACM International Conference on Information and Knowledge Management , ser. CIKM ’19, 2019, pp. 319–328. [Online]. Available: https://doi.org/10.1145/3357384.3358045
2019
Cited alongside, same era.
S. Lym, D. Lee, M. O’Connor, N. Chatterjee, and M. Erez, “DeLTA: GPU performance model for deep learning applications with in-depth memory system traffic analysis,” in 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , Mar. 2019, pp. 293–303
2019
Cited alongside, same era.
Z. Pei, C. Li, X. Qin, X. Chen, and G. Wei, “Iteration time prediction for CNN in multi-GPU platform: Modeling and analysis,” IEEE Access , vol. 7, pp. 64 788–64 797, 14 May 2019. [Online]. Available: https://doi.org/10.1109/ACCESS.2019.2916550
2019
Cited alongside, same era.
Q. Song, D. Cheng, H. Zhou, J. Yang, Y. Tian, and X. Hu, “Towards automated neural interaction discovery for click-through rate prediction,” in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , ser. KDD ’20, 2020, pp. 945–955. [Online]. Available: https://doi.org/10.1145/3394486.3403137
2020
Later among the works it cites.
W. Zhao, D. Xie, R. Jia, Y. Qian, R. Ding, M. Sun, and P. Li, “Distributed hierarchical GPU parameter server for massive scale deep learning ads systems,” in Proceedings of Machine Learning and Systems 2020 , ser. MLSys 2020, I. S. Dhillon, D. S. Papailiopoulos, and V. Sze, Eds. mlsys.org, Mar. 2020. [Online]. Available: https://proceedings.mlsys.org/book/315.pdf
2020
Later among the works it cites.
NVIDIA. (2020, Jul.) cuBLAS deep learning performance matrix multiplication. [Online]. Available: https://docs.nvidia.com/deeplearning/performance/dl-performance-matrix-multiplication/index.html
2020
Later among the works it cites.
Y.-C. Liao, C.-C. Wang, C.-H. Tu, M.-C. Kao, W.-Y. Liang, and S.-H. Hung, “PerfNetRT: Platform-aware performance modeling for optimized deep neural networks,” in 2020 International Computer Symposium , ser. ICS 2020, Dec. 2020, pp. 153–158. [Online]. Available: https://doi.org/10.1109/ICS51289.2020.00039
2020
Later among the works it cites.
H. Zhu, A. Phanishayee, and G. Pekhimenko, “Daydream: Accurately estimating the efficacy of optimizations for DNN training,” in 2020 USENIX Annual Technical Conference, USENIX ATC 2020, July 15-17, 2020 , A. Gavrilovska and E. Zadok, Eds. USENIX Association, 2020, pp. 337–352
2020
Later among the works it cites.
S. Li, R. J. Walls, and T. Guo, “Characterizing and modeling distributed training with transient cloud GPU servers,” in 2020 IEEE 40th International Conference on Distributed Computing Systems , ser. ICDCS 2020, Nov. 2020, pp. 943–953. [Online]. Available: https://doi.org/10.1109/ICDCS47774.2020.00097
2020
Later among the works it cites.
A. Tulloch. (2020, May) Batch embedding lookup GPU kernel and more. [Online]. Available: https://github.com/ajtulloch/sparse-ads-baselines
2020
Later among the works it cites.
2021
Later among the works it cites.
A. Rajagopal and C. Bouganis, “perf4sight: A toolflow to model cnn training performance on edge gpus,” in 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) . Los Alamitos, CA, USA: IEEE Computer Society, oct 2021, pp. 963–971. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/ICCVW54120.2021.00112
2021
Later among the works it cites.