Fetching the paper…
Reading the bibliography…
The use of GPUs has proliferated for machine learning workflows and is now considered mainstream for many deep learning models.
S. Williams, A. Waterman, and D. Patterson, “Roofline: an insightful visual performance model for multicore architectures,” Communications of the ACM , vol. 52, no. 4, pp. 65–76, 2009
2009
Earlier work this paper cites.
A. Thusoo, Z. Shao, S. Anthony, D. Borthakur, N. Jain, J. Sen Sarma, R. Murthy, and H. Liu, “Data warehousing and analytics infrastructure at Facebook,” in ACM International Conference on Management of data , 2010, pp. 1013–1020
2010
Earlier work this paper cites.
B. Recht, C. Re, S. Wright, and F. Niu, “Hogwild: A lock-free approach to parallelizing stochastic gradient descent,” in Advances in Neural Information Processing Systems 24 , J. Shawe-Taylor, R. S. Zemel, P. L. Bartlett, F. Pereira, and K. Q. Weinberger, Eds., 2011, pp. 693–701
2011
Earlier work this paper cites.
J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Y. Ng, “Large scale distributed deep networks,” in International Conference on Neural Information Processing Systems - Volume 1 , 2012, p. 1223–1231
2012
Earlier work this paper cites.
J. Dean and L. A. Barroso, “The tail at scale,” Communications of the ACM , vol. 56, no. 2, pp. 74–80, 2013
2013
Earlier work this paper cites.
I. Baldini, S. J. Fink, and E. Altman, “Predicting GPU performance from CPU runs using machine learning,” in IEEE International Symposium on Computer Architecture and High Performance Computing , 2014, pp. 254–261
2014
Earlier work this paper cites.
N. Ardalani, C. Lestourgeon, K. Sankaralingam, and X. Zhu, “Cross-architecture performance prediction (XAPP) using CPU code to predict GPU performance,” in International Symposium on Microarchitecture , 2015, pp. 725–737
2015
Earlier work this paper cites.
A. M. Elkahky, Y. Song, and X. He, “A multi-view deep learning approach for cross domain user modeling in recommendation systems,” in International Conference on World Wide Web , 2015, p. 278–288
2015
Earlier work this paper cites.
S. Zhang, A. E. Choromanska, and Y. LeCun, “Deep learning with elastic averaging sgd,” in Advances in neural information processing systems , 2015, pp. 685–693
2015
Earlier work this paper cites.
H.-T. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir, R. Anil, Z. Haque, L. Hong, V. Jain, X. Liu, and H. Shah, “Wide & deep learning for recommender systems,” in Workshop on Deep Learning for Recommender Systems , 2016
2016
Earlier work this paper cites.
P. Covington, J. Adams, and E. Sargin, “Deep neural networks for YouTube recommendations,” in ACM Recommender Systems Conference , 2016
2016
Earlier work this paper cites.
J. Dunn, “Introducing fblearner flow: Facebook’s AI backbone,” 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Dean, “Machine learning for systems and systems for machine learning,” in Presentation at the Conference on Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollar, and R. Girshick, “Mask R-CNN,” in IEEE International Conference on Computer Vision , Oct 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
“NVIDIA DGX,” https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/dgx-1/dgx-1-print-infographic-738238-nvidia-web.pdf , 2018
2018
Cited alongside, same era.
A. Qiao, B. Aragam, B. Zhang, and E. Xing, “Fault tolerance in iterative-convergent machine learning,” in International Conference on Machine Learning , 2019, pp. 5220–5230
2019
Later among the works it cites.
2019
Later among the works it cites.
C.-J. Wu, D. Brooks, U. Gupta, H.-H. Lee, and K. Hazelwood, “Deep learning: It’s not all about recognizing cats and dogs,” https://www.sigarch.org/deep-learning-its-not-all-about-recognizing-cats-and-dogs/ , 2019
2019
Later among the works it cites.
G. Zhou, N. Mou, Y. Fan, Q. Pi, W. Bian, C. Zhou, X. Zhu, and K. Gai, “Deep interest evolution network for click-through rate prediction,” in AAAI conference on artificial intelligence , vol. 33, 2019, pp. 5941–5948
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
K. Hazelwood, S. Bird, D. Brooks, S. Chintala, U. Diril, D. Dzhulgakov, M. Fawzy, B. Jia, Y. Jia, A. Kalro, J. Law, K. Lee, J. Lu, P. Noordhuis, M. Smelyanskiy, L. Xiong, and X. Wang, “Applied machine learning at Facebook: A datacenter infrastructure perspective,” in IEEE International Symposium on High Performance Computer Architecture , Feb 2018, pp. 620–629
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Y. Raimond and J. Basilico, “Deep learning for recommender systems,” https://www.slideshare.net/moustaki/deep-learning-for-recommender-systems-86752234 , January 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
X. Yi, Y.-F. Chen, S. Ramesh, V. Rajashekhar, L. Hong, N. Fiedel, N. Seshadri, L. Heldt, X. Wu, and E. H. Chi, “Factorized deep retrieval and distributed TensorFlow serving,” ser. Conference on Machine Learning and Systems, 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
I. Goldwasser, “Optimizing NVIDIA AI performance for MLPerf v0.7 training,” https://developer.nvidia.com/blog/optimizing-ai-performance-for-mlperf-v0-7-training/ , 2020, last accessed July 29, 2020
2020
Closest in time.
U. Gupta, S. Hsia, V. Saraph, X. Wang, B. Reagen, G.-Y. Wei, H.-H. S. Lee, D. Brooks, and C.-J. Wu, “DeepRecSys: A system for optimizing end-to-end at-scale neural recommendation inference,” 2020
2020
Closest in time.
U. Gupta, C.-J. Wu, X. Wang, M. Naumov, B. Reagen, D. Brooks, B. Cottel, K. Hazelwood, M. Hempstead, B. Jia et al. , “The architectural implications of Facebook’s DNN-based personalized recommendation,” in International Symposium on High Performance Computer Architecture , 2020, pp. 488–501
2020
Closest in time.
S. Hsia, U. Gupta, M. Wilkening, C.-J. Wu, G.-Y. Wei, and D. Brooks, “Cross-stack workload characterization of deep recommendation systems,” in IEEE International Symposium on Workload Characterization , 2020
2020
Closest in time.
L. Ke, U. Gupta, B. Y. Cho, D. Brooks, V. Chandra, U. Diril, A. Firoozshahian, K. Hazelwood, B. Jia, H.-H. S. Lee et al. , “RecNMP: Accelerating personalized recommendation with near-memory processing,” in ACM/IEEE International Symposium on Computer Architecture , 2020, pp. 790–803
2020
Closest in time.
K. Lee, “Introducing big basin: Our next-generation ai hardware,” https://fb.me/lee_2017 , 2017, last accessed April 17, 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
P. Mattson, V. J. Reddi, C. Cheng, C. Coleman, G. Diamos, D. Kanter, P. Micikevicius, D. Patterson, G. Schmuelling, H. Tang et al. , “MLPerf: An industry standard benchmark suite for machine learning performance,” IEEE Micro , vol. 40, no. 2, pp. 8–16, 2020
2020
Closest in time.
2020
Closest in time.
V. Nguyen, E. Oldridge, and M. Lee, “Announcing NVIDIA Merlin: An application framework for deep recommender systems,” https://developer.nvidia.com/blog/announcing-nvidia-merlin-application-framework-for-deep-recommender-systems/ , 2020, last accessed May 14, 2020
2020
Closest in time.
B. Nicolae, J. Li, J. M. Wozniak, G. Bosilca, M. Dorier, and F. Cappello, “DeepFreeze: Towards Scalable Asynchronous Checkpointing of Deep Learning Models,” in IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing , 2020. [Online]. Available: https://hal.inria.fr/hal-02543977
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.