Fetching the paper…
Reading the bibliography…
Recommender systems have been widely used in various large-scale user-oriented platforms for many years.
Factorization machines
S. Rendle · 2010
Earlier work this paper cites.
Introduction to recommender systems handbook
F. Ricci, L. Rokach, and B. Shapira · 2010
Earlier work this paper cites.
The word entropy of natural languages
C. Bentz and D. Alikaniotis · 2016
Earlier work this paper cites.
Wide & deep learning for recommender systems
H.-T. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir, et al · 2016
Earlier work this paper cites.
Deepfm: a factorization-machine based neural network for ctr prediction
H. Guo, R. Tang, Y. Ye, Z. Li, and X. He · 2017
Earlier work this paper cites.
Deep interest network for click-through rate prediction
G. Zhou, X. Zhu, C. Song, Y. Fan, H. Zhu, X. Ma, Y. Yan, J. Jin, H. Li, and K. Gai · 2018
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Search-based user interest modeling with lifelong sequential behavior data for click-through rate prediction
Q. Pi, G. Zhou, Y. Zhang, Z. Wang, L. Ren, Y. Fan, X. Zhu, and K. Gai · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Cited alongside, same era.
Large scale product graph construction for recommendation in e-commerce
X. Yang, Y. Zhu, Y. Zhang, X. Wang, and Q. Yuan · 2020
Cited alongside, same era.
A review of sparse expert models in deep learning
W. Fedus, J. Dean, and B. Zoph · 2022
Cited alongside, same era.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Cited alongside, same era.
Autoregressive image generation using residual quantization
D. Lee, C. Kim, S. Kim, M. Cho, and W.-S. Han · 2022
Minicpm: Unveiling the potential of small language models with scalable training strategies
S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, et al · 2024
Later among the works it cites.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al · 2024
Later among the works it cites.
A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al · 2024
Later among the works it cites.
Qarm: Quantitative alignment multi-modal recommendation at kuaishou
X. Luo, J. Cao, T. Sun, J. Yu, R. Huang, W. Yuan, H. Lin, Y. Zheng, S. Wang, Q. Hu, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Lamda: Language models for dialog applications
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, et al · 2022
Cited alongside, same era.
Twin: Two-stage interest network for lifelong user behavior modeling in ctr prediction at kuaishou
J. Chang, C. Zhang, Z. Fu, X. Zang, L. Guan, J. Lu, Y. Hui, D. Leng, Y. Niu, Y. Song, et al · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2023
Cited alongside, same era.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Cited alongside, same era.
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al · 2024
Cited alongside, same era.
S. Rajput, N. Mehta, A. Singh, R. Hulikal Keshavan, T. Vu, L. Heldt, L. Hong, Y. Tay, V. Tran, J. Samost, et al · 2024
Later among the works it cites.
Learning dynamics of llm finetuning
Y. Ren and D. J. Sutherland · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al · 2024
Later among the works it cites.
Twin v2: Scaling ultra-long user behavior sequence modeling for enhanced ctr prediction at kuaishou
Z. Si, L. Guan, Z. Sun, X. Zang, J. Lu, Y. Hui, X. Cao, Z. Yang, Y. Zheng, D. Leng, et al · 2024
Later among the works it cites.
Home: Hierarchy of multi-gate experts for multi-task learning at kuaishou
X. Wang, J. Cao, Z. Fu, K. Gai, and G. Zhou · 2024
Later among the works it cites.
Adapting large language models by integrating collaborative semantics for recommendation
B. Zheng, Y. Hou, H. Lu, Y. Chen, W. X. Zhao, M. Chen, and J.-R. Wen · 2024
Later among the works it cites.
Scaling the codebook size of vqgan to 100,000 with a utilization rate of 99%
L. Zhu, F. Wei, Y. Lu, and D. Chen · 2024
Later among the works it cites.
Pantheon: Personalized multi-objective ensemble sort via iterative pareto policy optimization
J. Cao, P. Xu, Y. Cheng, K. Guo, J. Tang, S. Wang, D. Leng, S. Yang, Z. Liu, Y. Niu, et al · 2025
Closest in time.