Fetching the paper…
Reading the bibliography…
In today's era of "scale-out", this paper makes the case that a specialized hardware architecture based on "scale-in"--placing as many specialized processors as possible along with their memory systems and interconnect links within one or two boards in a rack--would offer the potential to boost large recommender system throughput by 12-62x for inference and 12-45x for training compared to the DGX-2 state-of-the-art AI platform, while minimizing the performance impact of distributing large models across multiple processors.
1903
Earlier work this paper cites.
1906
Earlier work this paper cites.
1906
Earlier work this paper cites.
1908
Earlier work this paper cites.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
1912
Earlier work this paper cites.
2001
Earlier work this paper cites.
2003
Earlier work this paper cites.
2003
Earlier work this paper cites.
2005
Earlier work this paper cites.
2005
Earlier work this paper cites.
Chan, Ernie, Marcel Heimlich, Avi Purkayastha, and Robert Van De Geijn. “Collective communication: theory, practice, and experience.” Concurrency and Computation: Practice and Experience 19, no. 13 (2007): 1749-1783
2007
Earlier work this paper cites.
Varian, Hal R. “Online Ad auctions.” American Economic Review 99, no. 2 (2009): 430-34
2009
Earlier work this paper cites.
Rendle, Steffen. “Factorization machines.” In 2010 IEEE International Conference on Data Mining, pp. 995-1000. IEEE, 2010
2010
Earlier work this paper cites.
Graepel, Thore, Joaquin Quinonero Candela, Thomas Borchert, and Ralf Herbrich. “Web-scale bayesian click-through rate prediction for sponsored search advertising in microsoftś bing search engine.” Omnipress, 2010
2010
Cited alongside, same era.
He, Xinran, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi et al. “Practical lessons from predicting clicks on ads at facebook.” In Proceedings of the Eighth International Workshop on Data Mining for Online Advertising, pp. 1-9. 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
Li, Cheng, Yue Lu, Qiaozhu Mei, Dong Wang, and Sandeep Pandey. “Click-through prediction for advertising in twitter timeline.” In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1959-1968. 2015
“Zion: Facebook Next-Generation Large Memory Training Platform”, Misha Smelyanskiy, Hot Chips 31, August 19, 2019. See https://www.hotchips.org/hc31/HC31_1.10_MethodologyAndMLSystem-Facebook-rev-d.pdf
2019
Later among the works it cites.
Zhao, Weijie, Jingyuan Zhang, Deping Xie, Yulei Qian, Ronglai Jia, and Ping Li. “AIBox: CTR prediction model training on a single node.” In Proceedings of the 28th ACM International Conference on Information and Knowledge Management, pp. 319-328. 2019. See http://research.baidu.com/Public/uploads/5e18a1017a7a0.pdf
2019
Later among the works it cites.
Kwon, Youngeun, Yunjae Lee, and Minsoo Rhu. “Tensordimm: A practical near-memory processing architecture for embeddings and tensor operations in deep learning.” In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture, pp. 740-753. 2019
2019
Later among the works it cites.
Eitan Medina, Hot Chips 2019 presentation, Habana
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
Kim, Yoongu, Weikun Yang, and Onur Mutlu. “Ramulator: A fast and extensible DRAM simulator.” IEEE Computer architecture letters 15, no. 1 (2015): 45-49
2015
Cited alongside, same era.
Gomez-Uribe, Carlos A., and Neil Hunt. “The netflix recommender system: Algorithms, business value, and innovation.” ACM Transactions on Management Information Systems (TMIS) 6, no. 4 (2015): 1-19
2015
Cited alongside, same era.
Arunkumar, Akhil, Evgeny Bolotin, Benjamin Cho, Ugljesa Milic, Eiman Ebrahimi, Oreste Villa, Aamer Jaleel, Carole-Jean Wu, and David Nellans. “MCM-GPU: Multi-chip-module GPUs for continued performance scalability.” ACM SIGARCH Computer Architecture News 45, no. 2 (2017): 320-332
2017
Cited alongside, same era.
Peng, Ivy Bo, Roberto Gioiosa, Gokcen Kestor, Pietro Cicotti, Erwin Laure, and Stefano Markidis. “Exploring the performance benefit of hybrid memory system on HPC environments.” In 2017 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), pp. 683-692. IEEE, 2017
2017
Cited alongside, same era.
2018
Cited alongside, same era.
Hazelwood, Kim, Sarah Bird, David Brooks, Soumith Chintala, Utku Diril, Dmytro Dzhulgakov, Mohamed Fawzy et al. “Applied machine learning at facebook: A datacenter infrastructure perspective.” In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA), pp. 620-629. IEEE, 2018
2018
Cited alongside, same era.
Hadash, Guy, Oren Sar Shalom, and Rita Osadchy. “Rank and rate: multi-task learning for recommender systems.”, In Proceedings of the 12th ACM Conference on Recommender Systems, pp. 451-454. 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
“The true Processing In Memory accelerator”, Fabrice Devaux, Hot Chips 31, August 19, 2019. See https://www.hotchips.org/hc31/HC31_1.4_UPMEM.FabriceDevaux.v2_1.pdf
2019
Later among the works it cites.
Zhou, Guorui, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. “Deep interest evolution network for click-through rate prediction.” In Proceedings of the AAAI conference on artificial intelligence, vol. 33, pp. 5941-5948. 2019
2019
Later among the works it cites.
Zhao, Zhe, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, and Ed Chi. “Recommending what video to watch next: a multitask ranking system.” In Proceedings of the 13th ACM Conference on Recommender Systems, pp. 43-51. 2019
2019
Later among the works it cites.
https://www.nextplatform.com/2019/06/11/the-games-a-foot-intel-finally-gets-serious-about-ethernet-switching/
2019
Later among the works it cites.
Anil, Rohan, Vineet Gupta, Tomer Koren, and Yoram Singer. “Memory Efficient Adaptive Optimization.” In Advances in Neural Information Processing Systems, pp. 9749-9758. 2019
2019
Later among the works it cites.
“FPGA-based computing in the Era of Artificial Intelligence and Big Data”, Talk by Eriko Nurvitadhi, Intel Labs, ISPD 2019, http://ispd.cc/slides/2019/6_FPGASpecial_Eriko.pdf
2019
Later among the works it cites.
“DNN Accelerator Architectures”, Joel Emer, Vivienne Sze, Yu-Hsin Chen, ISCA Tutorial (2019), http://www.rle.mit.edu/eems/wp-content/uploads/2019/06/Tutorial-on-DNN-06-RS-Dataflow-and-NoC.pdf
2019
Later among the works it cites.
Ke, Liu, Udit Gupta, Benjamin Youngjae Cho, David Brooks, Vikas Chandra, Utku Diril, Amin Firoozshahian et al. “Recnmp: Accelerating personalized recommendation with near-memory processing.” In 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), pp. 790-803. IEEE, 2020
2020
Closest in time.
Li, Zhengjie, Yufan Zhang, Jian Wang, and Jinmei Lai. “A survey of FPGA design for AI era.” Journal of Semiconductors 41, no. 2 (2020): 021402
2020
Closest in time.
Chen, Yiran, Yuan Xie, Linghao Song, Fan Chen, and Tianqi Tang. “A Survey of Accelerator Architectures for Deep Neural Networks.” Engineering 6, no. 3 (2020): 264-274
2020
Closest in time.
Jouppi, Norman P., Doe Hyun Yoon, George Kurian, Sheng Li, Nishant Patil, James Laudon, Cliff Young, and David Patterson. “A domain-specific supercomputer for training deep neural networks.” Communications of the ACM 63, no. 7 (2020): 67-78
2020
Closest in time.