Fetching the paper…
Reading the bibliography…
In recommendation systems, practitioners observed that increase in the number of embedding tables and their sizes often leads to significant improvement in model performances.
Deep learning recommendation model for personalization and recommendation systems
Naumov, M., Mudigere, D., Shi, H. M., Huang, J., Sundaraman, N., Park, J., Wang, X., Gupta, U., Wu, C., Azzolini, A. G., Dzhulgakov, D., Mallevich, A., Cherniavskii, I., Lu, Y., Krishnamoorthi, R., Yu, A., Kondratenko, V., Pereira, S., Chen, X., Chen, W., Rao, V., Jia, B., Xiong, L., and Smelyanskiy, M · 1906
Earlier work this paper cites.
An analytical cache model
Agarwal, A. A., Hennessy, J. L., and Horowitz, M. H · 1989
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Practical lessons from predicting clicks on ads at facebook
He, X., Pan, J., Jin, O., Xu, T., Liu, B., Xu, T., Shi, Y., Atallah, A., Herbrich, R., Bowers, S., et al · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C · 2014
Earlier work this paper cites.
Andrews, M · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P · 2015
Earlier work this paper cites.
Topical word embeddings
Liu, Y., Liu, Z., Chua, T.-S., and Sun, M · 2015
Earlier work this paper cites.
Wide and deep learning for recommender systems
Cheng, H.-T., Koc, L., Harmsen, J., Shaked, T., Chandra, T., Aradhye, H., Anderson, G., Corrado, G., Chai, W., Ispir, M., Anil, R., Haque, Z., Hong, L., Jain, V., Liu, X., and Shah, H · 2016
Earlier work this paper cites.
How to train good word embeddings for biomedical nlp
Chiu, B., Crichton, G., Korhonen, A., and Pyysalo, S · 2016
Earlier work this paper cites.
Sparse word embeddings using l
Sun, F., Guo, J., Lan, Y., Xu, J., and Cheng, X · 2016
Earlier work this paper cites.
Compression-aware training of deep networks
Alvarez, J. M. and Salzmann, M · 2017
Earlier work this paper cites.
Fxpnet: Training a deep convolutional neural network in fixed-point representation
Chen, X., Hu, X., Zhou, H., and Xu, N · 2017
Cited alongside, same era.
Deep & cross network for ad click predictions
Wang, R., Fu, B., Fu, G., and Wang, M · 2017
Cited alongside, same era.
High-accuracy low-precision training
De Sa, C., Leszczynski, M., Zhang, J., Marzoev, A., Aberger, C. R., Olukotun, K., and Ré, C · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Cited alongside, same era.
Applied machine learning at Facebook: A datacenter infrastructure perspective
Hazelwood, K., Bird, S., Brooks, D., Chintala, S., Diril, U., Dzhulgakov, D., Fawzy, M., Jia, B., Jia, Y., Kalro, A., Law, J., Lee, K., Lu, J., Noordhuis, P., Smelyanskiy, M., Xiong, L., and Wang, X · 2018
Cited alongside, same era.
Block based singular value decomposition approach to matrix factorization for recommender systems
Bhavana, P., Kumar, V., and Padmanabhan, V · 2019
Later among the works it cites.
SinReQ: Generalized sinusoidal regularization for automatic low-bitwidth deep quantized training
Elthakeb, A. T., Pilligundla, P., and Esmaeilzadeh, H · 2019
Later among the works it cites.
Mixed dimension embeddings with application to memory-efficient recommendation systems
Ginart, A., Naumov, M., Mudigere, D., Yang, J. Y., and Zou, J · 2019
Later among the works it cites.
Post-training 4-bit quantization on embedding tables
Guan, H., Malevich, A., Yang, J., Park, J., and Yuen, H · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On periodic functions as regularizers for quantization of neural networks
Naumov, M., Diril, U., Park, J., Ray, B., Jablonski, J., and Tulloch, A · 2018
Cited alongside, same era.
Park, J., Naumov, M., Basu, P., Deng, S., Kalaiah, A., Khudia, D. S., Law, J., Malani, P., Malevich, A., Satish, N., Pino, J., Schatz, M., Sidorov, A., Sivakumar, V., Tulloch, A., Wang, X., Wu, Y., Yuen, H., Diril, U., Dzhulgakov, D., Hazelwood, K. M., Jia, B., Jia, Y., Qiao, L., Rao, V., Rotem, N., Yoo, S., and Smelyanskiy, M · 2018
Cited alongside, same era.
Dissecting contextual word embeddings: Architecture and representation
Peters, M., Neumann, M., Zettlemoyer, L., and Yih, W.-t · 2018
Cited alongside, same era.
Learning type-aware embeddings for fashion compatibility
Vasileva, M. I., Plummer, B. A., Dusad, K., Rajpal, S., Kumar, R., and Forsyth, D · 2018
Cited alongside, same era.
Fonduer: Knowledge base construction from richly formatted data
Wu, S., Hsiao, L., Cheng, X., Hancock, B., Rekatsinas, T., Levis, P., and Ré, C · 2018
Cited alongside, same era.
Training with low-precision embedding tables
Zhang, J., Yang, J., and Yuen, H · 2018
Cited alongside, same era.
Hierarchy-based image embeddings for semantic image retrieval
Barz, B. and Denzler, J · 2019
Cited alongside, same era.
Kalamkar, D., Mudigere, D., Mellempudi, N., Das, D., Banerjee, K., Avancha, S., Vooturi, D. T., Jammalamadaka, N., Huang, J., Yuen, H., Yang, J., Park, J., Heinecke, A., Georganas, E., Srinivasan, S., Kundu, A., Smelyanskiy, M., Kaul, B., and Dubey, P · 2019
Later among the works it cites.
Tensorized embedding layers for efficient model compression
Khrulkov, V., Hrinchuk, O., Mirvakhabova, L., and Oseledets, I. V · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Later among the works it cites.
Display Advertising Challenge
Criteo AI Lab · 2020
Closest in time.
Training with quantization noise for extreme model compression
Fan, A., Stock, P., Graham, B., Grave, E., Gribonval, R., Jégou, H., and Joulin, A · 2020
Closest in time.
Unified memory in CUDA 6
NVIDIA · 2020
Closest in time.
Distributed hierarchical gpu parameter server for massive scale deep learning ads systems
Zhao, W., Xie, D., Jia, R., Qian, Y., Ding, R., Sun, M., and Li, P · 2020
Closest in time.