Fetching the paper…
Reading the bibliography…
Scaling laws play an instrumental role in the sustainable improvement in model quality.
Factorization machines
Rendle, S · 2010
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Earlier work this paper cites.
Display advertising challenge, 2014
Tien, J.-B., joycenv, and Chapelle, O · 2014
Earlier work this paper cites.
The movielens datasets: History and context
Harper, F. M. and Konstan, J. A · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Higher-order factorization machines
Blondel, M., Fujino, A., Ueda, N., and Ishihata, M · 2016
Earlier work this paper cites.
Deep neural networks for youtube recommendations
Covington, P., Adams, J., and Sargin, E · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deepfm: a factorization-machine based neural network for ctr prediction
Guo, H., Tang, R., Ye, Y., Li, Z., and He, X · 2017
Earlier work this paper cites.
Motivating in-network aggregation for distributed deep neural network training
Luo, L., Liu, M., Nelson, J., Ceze, L., Phanishayee, A., and Krishnamurthy, A · 2017
Earlier work this paper cites.
Latent cross: Making use of context in recurrent recommender systems
Beutel, A., Covington, P., Jain, S., Xu, C., Li, J., Gatto, V., and Chi, E. H · 2018
Earlier work this paper cites.
Temporal hierarchical attention at category- and item-level for micro-video click-through prediction
Chen, X., Liu, D., Zha, Z.-J., Zhou, W., Xiong, Z., and Li, Y · 2018
Earlier work this paper cites.
xdeepfm: Combining explicit and implicit feature interactions for recommender systems
Lian, J., Zhou, X., Zhang, F., Chen, Z., Xie, X., and Sun, G · 2018
Earlier work this paper cites.
Parameter hub: a rack-scale parameter server for distributed deep neural network training
Luo, L., Nelson, J., Ceze, L., Phanishayee, A., and Krishnamurthy, A · 2018
Earlier work this paper cites.
Ad display/click data on taobao.com, 2018
Tianchi · 2018
Cited alongside, same era.
Dot product matrix compression for machine learning
Anonymous · 2019
Cited alongside, same era.
Deep learning recommendation model for personalization and recommendation systems
Naumov, M., Mudigere, D., Shi, H.-J. M., Huang, J., Sundaraman, N., Park, J., Wang, X., Gupta, U., Wu, C.-J., Azzolini, A. G., et al · 2019
Cited alongside, same era.
Autoint: Automatic feature interaction learning via self-attentive neural networks
Song, W., Shi, C., Xiao, Z., Duan, Z., Xu, Y., Zhang, M., and Tang, J · 2019
Cited alongside, same era.
Adaptive factorization network: Learning adaptive-order feature interactions
Cheng, W., Shen, Y., and Huang, L · 2020
Cited alongside, same era.
Learning to embed categorical features without embedding tables for recommendation
BARS: towards open benchmarking for recommender systems
Zhu, J., Dai, Q., Su, L., Ma, R., Liu, J., Cai, G., Xiao, X., and Zhang, R · 2022
Later among the works it cites.
BARS: towards open benchmarking for recommender systems
Zhu, J., Dai, Q., Su, L., Ma, R., Liu, J., Cai, G., Xiao, X., and Zhang, R · 2022
Later among the works it cites.
Vip5: Towards multimodal foundation models for recommendation
Geng, S., Tan, J., Liu, S., Fu, Z., and Zhang, Y · 2023
Later among the works it cites.
Hiformer: Heterogeneous feature interactions learning with transformers for recommender systems
Gui, H., Wang, R., Yin, K., Jin, L., Kula, M., Xu, T., Hong, L., and Chi, E. H · 2023
Later among the works it cites.
On the embedding collapse when scaling up recommendation models
Guo, X., Pan, J., Wang, X., Chen, B., Jiang, J., and Long, M · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kang, W.-C., Cheng, D. Z., Yao, T., Yi, X., Chen, T., Hong, L., and Chi, E. H · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Cited alongside, same era.
Persia: An open, hybrid system scaling deep learning-based recommenders up to 100 trillion parameters
Lian, X., Yuan, B., Zhu, X., Wang, Y., He, Y., Wu, H., Sun, L., Lyu, H., Liu, C., Dong, X., Liao, Y., Luo, M., Zhang, C., Xie, J., Li, H., Chen, L., Huang, R., Lin, J., Shu, C., Qiu, X., Liu, Z., Kong, D., Yuan, L., Yu, H., Yang, S., Zhang, C., and Liu, J · 2021
Cited alongside, same era.
High-performance, distributed training of large-scale deep learning recommendation models
Mudigere, D., Hao, Y., Huang, J., Tulloch, A., Sridharan, S., Liu, X., Ozdal, M., Nie, J., Park, J., Luo, L., et al · 2021
Cited alongside, same era.
Dcn v2: Improved deep & cross network and practical lessons for web-scale learning to rank systems
Wang, R., Shivanna, R., Cheng, D., Jain, S., Lin, D., Hong, L., and Chi, E · 2021
Cited alongside, same era.
Open benchmarking for click-through rate prediction
Zhu, J., Liu, J., Yang, S., Zhang, Q., and He, X · 2021
Cited alongside, same era.
Later among the works it cites.
How can recommender systems benefit from large language models: A survey
Lin, J., Dai, X., Xi, Y., Liu, W., Chen, B., Li, X., Zhu, C., Guo, H., Yu, Y., Tang, R., et al · 2023
Later among the works it cites.
Finalmlp: An enhanced two-stream mlp model for ctr prediction
Mao, K., Zhu, J., Su, L., Cai, G., Li, Y., and Dong, Z · 2023
Later among the works it cites.
Scaling law for recommendation models: Towards general-purpose user representations
Shin, K., Kwak, H., Kim, S. Y., Ramström, M. N., Jeong, J., Ha, J.-W., and Kim, K.-M · 2023
Later among the works it cites.
Improving training stability for multitask ranking models in recommender systems
Tang, J., Drori, Y., Chang, D., Sathiamoorthy, M., Gilmer, J., Wei, L., Yi, X., Hong, L., and Chi, E. H · 2023
Later among the works it cites.
Pre-train and search: Efficient embedding table sharding with pre-trained neural cost models
Zha, D., Feng, L., Luo, L., Bhushanam, B., Liu, Z., Hu, Y., Nie, J., Huang, Y., Tian, Y., Kejariwal, A., et al · 2023
Later among the works it cites.
Scaling law of large sequential recommendation models
Zhang, G., Hou, Y., Lu, H., Chen, Y., Zhao, W. X., and Wen, J.-R · 2023
Later among the works it cites.
Foundation models for recommender systems: A survey and new perspectives
Huang, C., Yu, T., Xie, K., Zhang, S., Yao, L., and McAuley, J · 2024
Closest in time.
Luo, L., Zhang, B., Tsang, M., Ma, Y., Chu, C.-H., Chen, Y., Li, S., Hao, Y., Zhao, Y., Lakshminarayanan, G., et al · 2024
Closest in time.
Feature fusion for the uninitiated | by siddharth sharma | medium
Sharma, S · 2024
Closest in time.