Fetching the paper…
Reading the bibliography…
The memory capacity of embedding tables in deep learning recommendation models (DLRMs) is increasing dramatically from tens of GBs to TBs across the industry.
Bayesian Tensorized Neural Networks with Automatic Rank Selection
Hawkins, C. and Zhang, Z · 1905
Earlier work this paper cites.
The Architectural Implications of Facebook’s DNN-based Personalized Recommendation
Gupta, U., Wu, C.-J., Wang, X., Naumov, M., Reagen, B., Brooks, D., Cottel, B., Hazelwood, K., Jia, B., Lee, H.-H. S., Malevich, A., Mudigere, D., Smelyanskiy, M., Xiong, L., and Zhang, X · 1906
Earlier work this paper cites.
Deep Learning Recommendation Model for Personalization and Recommendation Systems
Naumov, M., Mudigere, D., Shi, H.-J. M., Huang, J., Sundaraman, N., Park, J., Wang, X., Gupta, U., Wu, C.-J., Azzolini, A. G., Dzhulgakov, D., Mallevich, A., Cherniavskii, I., Lu, Y., Krishnamoorthi, R., Yu, A., Kondratenko, V., Pereira, S., Chen, X., Chen, W., Rao, V., Jia, B., Xiong, L., and Smelyanskiy, M · 1906
Earlier work this paper cites.
Principal component analysis
Wold, S., Esbensen, K., and Geladi, P · 1987
Earlier work this paper cites.
Feature hashing for large scale multitask learning
Weinberger, K., Dasgupta, A., Langford, J., Smola, A., and Attenberg, J · 2009
Earlier work this paper cites.
Tensor-train decomposition
Oseledets, I. V · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Restructuring of deep neural network acoustic models with singular value decomposition
Xue, J., Li, J., and Gong, Y · 2013
Earlier work this paper cites.
Tensorizing neural networks
Novikov, A., Podoprikhin, D., Osokin, A., and Vetrov, D · 2015
Earlier work this paper cites.
Toward accelerating deep learning at scale using specialized hardware in the datacenter
Ovtcharov, K., Ruwase, O., Kim, J.-Y., Fowers, J., Strauss, K., and Chung, E · 2015
Earlier work this paper cites.
On the expressive power of deep learning: A tensor analysis
Cohen, N., Sharir, O., and Shashua, A · 2016
Earlier work this paper cites.
Gpu kernels for block-sparse weights
Gray, S., Radford, A., and Kingma, D. P · 2017
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., Boyle, R., Cantin, P., Chao, C., Clark, C., Coriell, J., Daley, M., Dau, M., Dean, J., Gelb, B., Ghaemmaghami, T. V., Gottipati, R., Gulland, W., Hagmann, R., Ho, C. R., Hogberg, D., Hu, J., Hundt, R., Hurt, D., Ibarz, J., Jaffey, A., Jaworski, A., Kaplan, A., Khaitan, H., Killebrew, D., Koch, A., Kumar, N., Lacy, S., Laudon, J., Law, J., Le, D., Leary, C., Liu, Z., Lucke, K., Lundin, A., MacKean, G., Maggiore, A., Mahony, M., Miller, K., Nagarajan, R., Narayanaswami, R., Ni, R., Nix, K., Norrie, T., Omernick, M., Penukonda, N., Phelps, A., Ross, J., Ross, M., Salek, A., Samadiani, E., Severn, C., Sizikov, G., Snelham, M., Souter, J., Steinberg, D., Swing, A., Tan, M., Thorson, G., Tian, B., Toma, H., Tuttle, E., Vasudevan, V., Walter, R., Wang, W., Wilcox, E., and Yoon, D. H · 2017
Earlier work this paper cites.
Learning sparse neural networks through l _ 0 l\_0 regularization
Louizos, C., Welling, M., and Kingma, D. P · 2017
Earlier work this paper cites.
Tensor-train recurrent neural networks for video classification
Yang, Y., Krompass, D., and Tresp, V · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Cited alongside, same era.
Serving dnns in real time at datacenter scale with project brainwave
Chung, E., Fowers, J., Ovtcharov, K., Papamichael, M., Caulfield, A., Massengill, T., Liu, M., Ghandi, M., Lo, D., Reinhardt, S., Alkalay, S., Angepat, H., Chiou, D., Forin, A., Burger, D., Woods, L., Weisz, G., Haselman, M., and Zhang, D · 2018
Cited alongside, same era.
A configurable cloud-scale dnn processor for real-time ai
Fowers, J., Ovtcharov, K., Papamichael, M., Massengill, T., Liu, M., Lo, D., Alkalay, S., Haselman, M., Adams, L., Ghandi, M., Heil, S., Patel, P., Sapek, A., Weisz, G., Woods, L., Lanka, S., Reinhardt, S., Caulfield, A., Chung, E., and Burger, D · 2018
Cited alongside, same era.
Applied machine learning at facebook: A datacenter infrastructure perspective
Hazelwood, K., Bird, S., Brooks, D., Chintala, S., Diril, U., Dzhulgakov, D., Fawzy, M., Jia, B., Jia, Y., Kalro, A., Law, J., Lee, K., Lu, J., Noordhuis, P., Smelyanskiy, M., Xiong, L., and Wang, X · 2018
Reddi, V. J., Cheng, C., Kanter, D., Mattson, P., Schmuelling, G., Wu, C.-J., Anderson, B., Breughe, M., Charlebois, M., Chou, W., Chukka, R., Coleman, C., Davis, S., Deng, P., Diamos, G., Duke, J., Fick, D., Gardner, J. S., Hubara, I., Idgunji, S., Jablin, T. B., Jiao, J., John, T. S., Kanwar, P., Lee, D., Liao, J., Lokhmotov, A., Massa, F., Meng, P., Micikevicius, P., Osborne, C., Pekhimenko, G., Rajan, A. T. R., Sequeira, D., Sirasao, A., Sun, F., Tang, H., Thomson, M., Wei, F., Wu, E., Xu, L., Yamada, K., Yu, B., Yuan, G., Zhong, A., Zhang, P., and Zhou, Y · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
Training with multi-layer embeddings for model reduction
Ghaemmaghami, B., Deng, Z., Cho, B., Orshansky, L., Singh, A. K., Erez, M., and Orshansky, M · 2020
Later among the works it cites.
Deeprecsys: A system for optimizing end-to-end at-scale neural recommendation inference
Gupta, U., Hsia, S., Saraph, V., Wang, X., Reagen, B., Wei, G., Lee, H. S., Brooks, D., and Wu, C · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Wide Compression: Tensor Ring Nets
Wang, W., Sun, Y., Eriksson, B., Wang, W., and Aggarwal, V · 2018
Cited alongside, same era.
Alibaba unveils ai chip to enhance cloud computing power
Alibaba · 2019
Cited alongside, same era.
AWS Inferentia: High performance machine learning inference chip, custom designed by AWS
Amazon · 2019
Cited alongside, same era.
Cloud TPU: Codesigning architecture and infrastructure, 2019
Chao, C. and Saeta, B · 2019
Cited alongside, same era.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Cited alongside, same era.
Post-training 4-bit quantization on embedding tables, 2019
Guan, H., Malevich, A., Yang, J., Park, J., and Yuen, H · 2019
Cited alongside, same era.
Accelerating facebook’s infrastructure with application-specific hardware
Lee, K. and Rao, V · 2019
Cited alongside, same era.
Later among the works it cites.
Deep learning: It’s not all about recognizing cats and dogs, 2020
Hazelwood, K · 2020
Later among the works it cites.
Tensorized embedding layers for efficient model compression, 2020
Hrinchuk, O., Khrulkov, V., Mirvakhabova, L., Orlova, E., and Oseledets, I · 2020
Later among the works it cites.
Recnmp: Accelerating personalized recommendation with near-memory processing
Ke, L., Gupta, U., Cho, B. Y., Brooks, D., Chandra, V., Diril, U., Firoozshahian, A., Hazelwood, K., Jia, B., Lee, H. S., Li, M., Maher, B., Mudigere, D., Naumov, M., Schatz, M., Smelyanskiy, M., Wang, X., Reagen, B., Wu, C., Hempstead, M., and Zhang, X · 2020
Later among the works it cites.
Mlperf: An industry standard benchmark suite for machine learning performance
Mattson, P., Reddi, V. J., Cheng, C., Coleman, C., Diamos, G., Kanter, D., Micikevicius, P., Patterson, D., Schmuelling, G., Tang, H., Wei, G.-w., and Wu, C.-J · 2020
Later among the works it cites.
MLPerf Training
MLPerf · 2020
Later among the works it cites.
Deep learning training in facebook data centers: Design of scale-up and scale-out systems, 2020
Naumov, M., Kim, J., Mudigere, D., Sridharan, S., Wang, X., Zhao, W., Yilmaz, S., Kim, C., Yuen, H., Ozdal, M., Nair, K., Gao, I., Su, B.-Y., Yang, J., and Smelyanskiy, M · 2020
Later among the works it cites.
2.8.13. cublasgemmbatchedex, cublas :: Cuda toolkit documentation, 2020
NVIDIA · 2020
Later among the works it cites.
Tensorized Embedding Layers for Efficient Model Compression
Rusu, A. A., Rao, D., Sygnowski, J., Vinyals, O., Pascanu, R., Osindero, S., and Hadsell, R · 2020
Later among the works it cites.
Developing a recommendation benchmark for mlperf training and inference, 2020
Wu, C.-J., Burke, R., Chi, E. H., Konstan, J., McAuley, J., Raimond, Y., and Zhang, H · 2020
Later among the works it cites.
Distributed hierarchical gpu parameter server for massive scale deep learning ads systems
Zhao, W., Xie, D., Jia, R., Qian, Y., Ding, R., Sun, M., and Li, P · 2020
Later among the works it cites.