Fetching the paper…
Reading the bibliography…
Learning high-quality feature embeddings efficiently and effectively is critical for the performance of web-scale machine learning systems.
Five Lessons: The Modern Fundamentals of Golf
B. Hogan and H. W. Wind · 1985
Earlier work this paper cites.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Fast learning in multi-resolution hierarchies
J. Moody · 1988
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, and P. Vincent · 2000
Earlier work this paper cites.
The power of two choices in randomized load balancing
M. Mitzenmacher · 2001
Earlier work this paper cites.
Database-friendly random projections: Johnson-Lindenstrauss with binary coins
D. Achlioptas · 2003
Earlier work this paper cites.
An elementary proof of a theorem of Johnson and Lindenstrauss
S. Dasgupta and A. Gupta · 2003
Earlier work this paper cites.
Very sparse random projections
P. Li, T. J. Hastie, and K. W. Church · 2006
Earlier work this paper cites.
Feature hashing for large scale multitask learning
K. Weinberger, A. Dasgupta, J. Langford, A. Smola, and J. Attenberg · 2009
Earlier work this paper cites.
Pseudorandomness
S. P. Vadhan et al · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Modeling delayed feedback in display advertising
O. Chapelle · 2014
Earlier work this paper cites.
Compressing neural networks with the hashing trick
W. Chen, J. Wilson, S. Tyree, K. Weinberger, and Y. Chen · 2015
Earlier work this paper cites.
The movielens datasets: History and context
F. M. Harper and J. A. Konstan · 2015
Earlier work this paper cites.
Deep neural networks for YouTube recommendations
P. Covington, J. Adams, and E. Sargin · 2016
Cited alongside, same era.
Practical hash functions for similarity estimation and dimensionality reduction
S. Dahlgaard, M. Knudsen, and M. Thorup · 2017
Cited alongside, same era.
Hash embeddings for efficient word representations
D. Svenstrup, J. Hansen, and O. Winther · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Fully understanding the hashing trick
C. B. Freksen, L. Kamma, and K. Green Larsen · 2018
Cited alongside, same era.
Deep learning recommendation model for personalization and recommendation systems
M. Naumov, D. Mudigere, H.-J. M. Shi, J. Huang, N. Sundaraman, J. Park, X. Wang, U. Gupta, C.-J. Wu, A. G. Azzolini, et al · 2019
Learning to embed categorical features without embedding tables for recommendation
W.-C. Kang, D. Z. Cheng, T. Yao, X. Yi, T. Chen, L. Hong, and E. H. Chi · 2021
Later among the works it cites.
DCN V2: Improved deep & cross network and practical lessons for web-scale learning to rank systems
R. Wang, R. Shivanna, D. Z. Cheng, S. Jain, D. Lin, L. Hong, and E. H. Chi · 2021
Later among the works it cites.
On the factory floor: ML engineering for industrial-scale ads recommendation models
R. Anil, S. Gadanho, D. Huang, N. Jacob, Z. Li, D. Lin, T. Phillips, C. Pop, K. Regan, G. I. Shamir, R. Shivanna, and Q. Yan · 2022
Later among the works it cites.
The trade-offs of model size in large recommendation models : 100GB to 10MB Criteo-tb DLRM model
A. Desai and A. Shrivastava · 2022
Later among the works it cites.
Random offset block embedding (ROBE) for compressed embedding tables in deep learning recommendation systems
A. Desai, L. Chou, and A. Shrivastava · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Autoint: Automatic feature interaction learning via self-attentive neural networks
W. Song, C. Shi, Z. Xiao, Z. Duan, Y. Xu, M. Zhang, and J. Tang · 2019
Cited alongside, same era.
Sampling-bias-corrected neural modeling for large corpus item recommendations
X. Yi, J. Yang, L. Hong, D. Z. Cheng, L. Heldt, A. Kumthekar, Z. Zhao, L. Wei, and E. H. Chi · 2019
Cited alongside, same era.
Recommending what video to watch next: A multitask ranking system
Z. Zhao, L. Hong, L. Wei, J. Chen, A. Nath, S. Andrews, A. Kumthekar, M. Sathiamoorthy, X. Yi, and E. H. Chi · 2019
Cited alongside, same era.
Can weight sharing outperform random architecture search? An investigation with TuNAS
G. Bender, H. Liu, B. Chen, G. Chu, S. Cheng, P. Kindermans, and Q. V. Le · 2020
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
Prevalence of neural collapse during the terminal phase of deep learning training
V. Papyan, X. Han, and D. L. Donoho · 2020
Cited alongside, same era.
Software-hardware co-design for fast and scalable training of deep learning recommendation models
D. Mudigere, Y. Hao, J. Huang, Z. Jia, A. Tulloch, S. Sridharan, X. Liu, M. Ozdal, J. Nie, J. Park, L. Luo, J. A. Yang, L. Gao, D. Ivchenko, A. Basant, Y. Hu, J. Yang, E. K. Ardestani, X. Wang, R. Komuravelli, C. Chu, S. Yilmaz, H. Li, J. Qian, Z. Feng, Y. Ma, J. Yang, E. Wen, H. Li, L. Yang, C. Sun, W. Zhao, D. Melts, K. Dhulipala, K. R. Kishore, T. Graf, A. Eisenman, K. K. Matam, A. Gangidi, G. J. Chen, M. Krishnan, A. Nayak, K. Nair, B. Muthiah, M. khorashadi, P. Bhattacharya, P. Lapukhov, M. Naumov, A. Mathews, L. Qiao, M. Smelyanskiy, B. Jia, and V. Rao · 2022
Later among the works it cites.
Clustering embedding tables, without first learning them
H. L.-H. Tsang and T. D. Ahle · 2022
Later among the works it cites.
Bars: Towards open benchmarking for recommender systems
J. Zhu, Q. Dai, L. Su, R. Ma, J. Liu, G. Cai, X. Xiao, and R. Zhang · 2022
Later among the works it cites.
Performance of ℓ 1 \ell_{1} regularization for sparse convex optimization
K. Axiotis and T. Yasuda · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with GPT-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Closest in time.
Learning rate schedules in the presence of distribution shift
M. Fahrbach, A. Javanmard, V. Mirrokni, and P. Worah · 2023
Closest in time.
Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings
N. P. Jouppi, G. Kurian, S. Li, P. Ma, R. Nagarajan, L. Nai, N. Patil, S. Subramanian, A. Swing, B. Towles, et al · 2023
Closest in time.
Sequential attention for feature selection
T. Yasuda, M. Bateni, L. Chen, M. Fahrbach, G. Fu, and V. Mirrokni · 2023
Closest in time.