Fetching the paper…
Reading the bibliography…
While providing low latency is a fundamental requirement in deploying recommendation services, achieving high resource utility is also crucial in cost-effectively maintaining the datacenter.
Y. Jiang, K. Tian, and X. Shen, “Combining Locality Analysis with Online Proactive Job Co-Scheduling in Chip Multiprocessors,” in International Conference on High-Performance Embedded Architectures and Compilers
2010
Earlier work this paper cites.
Y. Jiang, K. Tian, X. Shen, J. Zhang, J. Chen, and R. Tripathi, “The Complexity of Optimal Job Co-Scheduling on Chip Multiprocessors and Heuristics-Based Solutions,” IEEE Transactions on Parallel and Distributed Systems
2010
Earlier work this paper cites.
J. Mars, L. Tang, R. Hundt, K. Skadron, and M. L. Soffa, “Bubble-Up: Increasing Utilization in Modern Warehouse Scale Computers via Sensible Co-Locations,” in Proceedings of the International Symposium on Microarchitecture (MICRO)
2011
Earlier work this paper cites.
C. Delimitrou and C. Kozyrakis, “Paragon: QoS-Aware Scheduling for Heterogeneous Datacenters,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS)
2013
Earlier work this paper cites.
H. Kasture and D. Sanchez, “Ubik: Efficient Cache Sharing with Strict QoS for Latency-Critical Workloads,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS)
2014
Earlier work this paper cites.
D. Lo, L. Cheng, R. Govindaraju, P. Ranganathan, and C. Kozyrakis, “Heracles: Improving Resource Efficiency at Scale,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2015
Earlier work this paper cites.
H.-T. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir, R. Anil, Z. Haque, L. Hong, V. Jain, X. Liu, and H. Shah, “Wide & Deep Learning for Recommender Systems,” in Proceedings of the 1st workshop on deep learning for recommender systems
2016
Earlier work this paper cites.
X. He, L. Liao, H. Zhang, L. Nie, X. Hu, and T.-S. Chua, “Neural Collaborative Filtering,” in Proceedings of the international conference on world wide web
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
G. Zhou, X. Zhu, C. Song, Y. Fan, H. Zhu, X. Ma, Y. Yan, J. Jin, H. Li, and K. Gai, “Deep Interest Network for Click-Through Rate Prediction,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
K. Hazelwood, S. Bird, D. Brooks, S. Chintala, U. Diril, D. Dzhulgakov, M. Fawzy, B. Jia, Y. Jia, A. Kalro, J. Law, K. Lee, J. Lu, P. Noordhuis, M. Smelyanskiy, L. Xiong, and X. Wang, “Applied Machine Learning at Facebook: A Datacenter Infrastructure Perspective,” in Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA)
2018
Earlier work this paper cites.
X. Yi, Y.-F. Chen, S. Ramesh, V. Rajashekhar, L. Hong, N. Fiedel, N. Seshadri, L. Heldt, X. Wu, and H. Chi, “Factorized Deep Retrieval and Distributed TensorFlow Serving,” in Conference on Machine Learning and Systems
2018
Earlier work this paper cites.
T. Singhal, “Maximizing GPU Utilization For Datacenter Inference with NVIDIA TensorRT Inference Server.” NVIDIA Webinar, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
G. Zhou, N. Mou, Y. Fan, Q. Pi, W. Bian, C. Zhou, X. Zhu, and K. Gai, “Deep Interest Evolution Network for Click-Through Rate Prediction,” in Proceedings of the AAAI Conference on Artificial Intelligence
2019
Cited alongside, same era.
S. Chen, C. Delimitrou, and J. F. Martínez, “PARTIES: Qos-Aware Resource Partitioning for Multiple Interactive Services,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS)
2019
Y. Choi and M. Rhu, “PREMA: A Predictive Multi-task Scheduling Algorithm For Preemptible Neural Processing Units,” in Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA)
2020
Later among the works it cites.
S. Ghodrati, B. H. Ahn, J. K. Kim, S. Kinzer, B. R. Yatham, N. Alla, H. Sharma, M. Alian, E. Ebrahimi, N. S. Kim, C. Young, and H. Esmaeilzadeh, “Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural Networks,” in Proceedings of the International Symposium on Microarchitecture (MICRO)
2020
Later among the works it cites.
L. Ke, U. Gupta, B. Y. Cho, D. Brooks, V. Chandra, U. Diril, A. Firoozshahian, K. Hazelwood, B. Jia, H.-H. S. Lee, M. Li, B. Maher, D. Mudigere, M. Naumov, M. Schatz, M. Smelyanskiy, X. Wang, B. Reagen, C.-J. Wu, M. Hempstead, and X. Zhang, “RecNMP: Accelerating Personalized Recommendation with Near-Memory Processing,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Y. Kwon, Y. Lee, and M. Rhu, “TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning,” in Proceedings of the International Symposium on Microarchitecture (MICRO)
2019
Cited alongside, same era.
W. Zhao, J. Zhang, D. Xie, Y. Qian, R. Jia, and P. Li, “AIBox: CTR Prediction Model Training on A Single Node,” in Proceedings of the ACM International Conference on Information and Knowledge Management
2019
Cited alongside, same era.
J. Hestness, N. Ardalani, and G. Diamos, “Beyond Human-Level Accuracy: Computational Challenges in Deep Learning,” in Proceedings of the Symposium on Principles and Practice of Parallel Programming (PPOPP)
2019
Cited alongside, same era.
U. Gupta, C.-J. Wu, X. Wang, M. Naumov, B. Reagen, D. Brooks, B. Cottel, K. Hazelwood, M. Hempstead, B. Jia, H.-H. S. Lee, A. Malevich, D. Mudigere, M. Smelyanskiy, L. Xiong, and X. Zhang, “The Architectural Implications of Facebook’s DNN-based Personalized Recommendation,” in Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA)
2020
Cited alongside, same era.
2020
Cited alongside, same era.
R. Hwang, T. Kim, Y. Kwon, and M. Rhu, “Centaur: A Chiplet-Based, Hybrid Sparse-Dense Accelerator for Personalized Recommendations,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2020
Cited alongside, same era.
U. Gupta, S. Hsia, V. Saraph, X. Wang, B. Reagen, G.-Y. Wei, H.-H. S. Lee, D. Brooks, and C.-J. Wu, “DeepRecSys: A System for Optimizing End-to-end At-scale Neural Recommendation Inference,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2020
Cited alongside, same era.
T. Patel and D. Tiwari, “CLITE: Efficient and Qos-Aware Co-location of Multiple Latency-Critical Jobs for Warehouse Scale Computers,” in Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA)
2020
Cited alongside, same era.
B. Kim, J. Park, E. Lee, M. Rhu, and J. H. Ahn, “TRiM: Tensor Reduction in Memory,” IEEE Computer Architecture Letters
2020
Later among the works it cites.
2020
Later among the works it cites.
SambaNova, “Surpassing State-of-the-Art Accuracy in Recommendation Models.” https://sambanova.ai/blog/surpassing-state-of-the-art-accuracy-in-recommendation-models/ , 2020
2020
Later among the works it cites.
W. Zhao, D. Xie, R. Jia, Y. Qian, R. Ding, M. Sun, and P. Li, “Distributed Hierarchical GPU Parameter Server for Massive Scale Deep Learning Ads Systems,” in Proceedings of Machine Learning and Systems
2020
Later among the works it cites.
Z. Deng, J. Park, P. T. P. Tang, H. Liu, J. Yang, H. Yuen, J. Huang, D. Khudia, X. Wei, E. Wen, D. Choudhary, R. Krishnamoorthi, C.-J. Wu, S. Nadathur, C. Kim, M. Naumov, S. Naghshineh, and M. Smelyanskiy, “Low-Precision Hardware Architectures Meet Recommendation Model Inference at Scale,” 2021
2021
Later among the works it cites.
U. Gupta, S. Hsia, J. Zhang, M. Wilkening, J. Pombra, H.-H. S. Lee, G.-Y. Wei, C.-J. Wu, and D. Brooks, “RecPipe: Co-Designing Models and Hardware to Jointly Optimize Recommendation Quality and Performance,” in Proceedings of the International Symposium on Microarchitecture (MICRO)
2021
Later among the works it cites.
H. Liu, Q. Gao, J. Li, X. Liao, H. Xiong, G. Chen, W. Wang, G. Yang, Z. Zha, D. Dong, D. Dou, and H. Xiong, “JIZHI: A Fast and Cost-Effective Model-As-A-Service System for Web-Scale Online Inference at Baidu,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
2021
Later among the works it cites.
Y. Kwon, Y. Lee, and M. Rhu, “Tensor Casting: Co-Designing Algorithm-Architecture for Personalized Recommendation Training,” in Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA)
2021
Later among the works it cites.
M. Wilkening, U. Gupta, S. Hsia, C. Trippel, C.-J. Wu, D. Brooks, and G.-Y. Wei, “RecSSD: Near Data Processing for Solid State Drive Based Recommendation Inference,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS)
2021
Later among the works it cites.