Fetching the paper…
Reading the bibliography…
Large-scale recommendation systems are characterized by their reliance on high cardinality, heterogeneous features and the need to handle tens of billions of user actions on a daily basis.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 1904
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Clustering for approximate similarity search in high-dimensional spaces
Li, C., Chang, E., Garcia-Molina, H., and Wiederhold, G · 2002
Earlier work this paper cites.
Product quantization for nearest neighbor search
Jegou, H., Douze, M., and Schmid, C · 2010
Earlier work this paper cites.
Factorization machines
Rendle, S · 2010
Earlier work this paper cites.
Atlas: A probabilistic algorithm for high dimensional similarity search
Zhai, J., Lou, Y., and Gehrke, J · 2011
Earlier work this paper cites.
Expectation-maximization for learning determinantal point processes
Gillenwater, J., Kulesza, A., Fox, E., and Taskar, B · 2014
Earlier work this paper cites.
Training highly multiclass classifiers
Gupta, M. R., Bengio, S., and Weston, J · 2014
Earlier work this paper cites.
Practical lessons from predicting clicks on ads at facebook
He, X., Pan, J., Jin, O., Xu, T., Liu, B., Xu, T., Shi, Y., Atallah, A., Herbrich, R., Bowers, S., and Candela, J. Q · 2014
Earlier work this paper cites.
Asymmetric lsh (alsh) for sublinear time maximum inner product search (mips)
Shrivastava, A. and Li, P · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Wide & deep learning for recommender systems
Cheng, H.-T., Koc, L., Harmsen, J., Shaked, T., Chandra, T., Aradhye, H., Anderson, G., Corrado, G., Chai, W., Ispir, M., Anil, R., Haque, Z., Hong, L., Jain, V., Liu, X., and Shah, H · 2016
Earlier work this paper cites.
Deep neural networks for youtube recommendations
Covington, P., Adams, J., and Sargin, E · 2016
Earlier work this paper cites.
Session-based recommendations with recurrent neural networks
Hidasi, B., Karatzoglou, A., Baltrunas, L., and Tikk, D · 2016
Earlier work this paper cites.
Deep networks with stochastic depth, 2016
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K · 2016
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Elfwing, S., Uchibe, E., and Doya, K · 2017
Earlier work this paper cites.
Deepfm: A factorization-machine based neural network for ctr prediction
Guo, H., Tang, R., Ye, Y., Li, Z., and He, X · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Attentional factorization machines: Learning the weight of feature interactions via attention networks
Xiao, J., Ye, H., He, X., Zhang, H., Wu, F., and Chua, T.-S · 2017
Earlier work this paper cites.
Pixie: A system for recommending 3+ billion items to 200+ million users in real-time
Eksombatchai, C., Jindal, P., Liu, J. Z., Liu, Y., Sharma, R., Sugnet, C., Ulrich, M., and Leskovec, J · 2018
Earlier work this paper cites.
Self-attentive sequential recommendation
Kang, W.-C. and McAuley, J · 2018
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Ma, J., Zhao, Z., Yi, X., Chen, J., Hong, L., and Chi, E. H · 2018
Earlier work this paper cites.
Deep reinforcement learning for page-wise recommendations
Zhao, X., Xia, L., Zhang, L., Ding, Z., Yin, D., and Tang, J · 2018
Earlier work this paper cites.
Deep interest network for click-through rate prediction
Zhou, G., Zhu, X., Song, C., Fan, Y., Zhu, H., Ma, X., Yan, Y., Jin, J., Li, H., and Gai, K · 2018
Cited alongside, same era.
Behavior sequence transformer for e-commerce recommendation in alibaba
Chen, Q., Zhao, H., Li, W., Huang, P., and Ou, W · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Bert4rec: Sequential recommendation with bidirectional encoder representations from transformer
Sun, F., Liu, J., Wu, J., Pei, C., Lin, X., Ou, W., and Jiang, P · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Later among the works it cites.
Efficiently modeling long sequences with structured state spaces
Gu, A., Goel, K., and Ré, C · 2022
Later among the works it cites.
Transformer quality in linear time
Hua, W., Dai, Z., Liu, H., and Le, Q. V · 2022
Later among the works it cites.
Reducing activation recomputation in large transformer models, 2022
Korthikanti, V., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B · 2022
Later among the works it cites.
Monolith: Real time recommendation system with collisionless embedding table, 2022
Liu, Z., Zou, L., Zou, X., Wang, C., Zhang, B., Tang, D., Zhu, B., Zhu, Y., Wu, P., Wang, K., and Cheng, Y · 2022
Later among the works it cites.
Software-hardware co-design for fast and scalable training of deep learning recommendation models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Neural collaborative filtering vs. matrix factorization revisited
Rendle, S., Krichene, W., Zhang, L., and Anderson, J · 2020
Cited alongside, same era.
Glu variants improve transformer, 2020
Shazeer, N · 2020
Cited alongside, same era.
Progressive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations
Tang, H., Liu, J., Zhao, M., and Gong, X · 2020
Cited alongside, same era.
Cold: Towards the next generation of pre-ranking system, 2020
Wang, Z., Zhao, L., Jiang, B., Zhou, G., Zhu, X., and Gai, K · 2020
Cited alongside, same era.
Mudigere, D., Hao, Y., Huang, J., Jia, Z., Tulloch, A., Sridharan, S., Liu, X., Ozdal, M., Nie, J., Park, J., Luo, L., Yang, J. A., Gao, L., Ivchenko, D., Basant, A., Hu, Y., Yang, J., Ardestani, E. K., Wang, X., Komuravelli, R., Chu, C.-H., Yilmaz, S., Li, H., Qian, J., Feng, Z., Ma, Y., Yang, J., Wen, E., Li, H., Yang, L., Sun, C., Zhao, W., Melts, D., Dhulipala, K., Kishore, K., Graf, T., Eisenman, A., Matam, K. K., Gangidi, A., Chen, G. J., Krishnan, M., Nayak, A., Nair, K., Muthiah, B., khorashadi, M., Bhattacharya, P., Lapukhov, P., Naumov, M., Mathews, A., Qiao, L., Smelyanskiy, M., Jia, B., and Rao, V · 2022
Later among the works it cites.
Efficiently scaling transformer inference, 2022
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Levskaya, A., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2022
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2022
Later among the works it cites.
Zero-shot recommendation as language modeling
Sileo, D., Vossen, W., and Raymaekers, R · 2022
Later among the works it cites.
Dhen: A deep and hierarchical ensemble network for large-scale click-through rate prediction, 2022
Zhang, B., Luo, L., Liu, X., Li, J., Chen, Z., Zhang, W., Wei, X., Hao, Y., Tsang, M., Wang, W., Liu, Y., Li, H., Badr, Y., Park, J., Yang, J., Mudigere, D., and Wen, E · 2022
Later among the works it cites.
Tallrec: An effective and efficient tuning framework to align large language model with recommendation
Bao, K., Zhang, J., Zhang, Y., Wang, W., Feng, F., and He, X · 2023
Later among the works it cites.
Twin: Two-stage interest network for lifelong user behavior modeling in ctr prediction at kuaishou, 2023
Chang, J., Zhang, C., Fu, Z., Zang, X., Guan, L., Lu, J., Hui, Y., Leng, D., Niu, Y., Song, Y., and Gai, K · 2023
Later among the works it cites.
Flashattention-2: Faster attention with better parallelism and work partitioning, 2023
Dao, T · 2023
Later among the works it cites.
Turning dross into gold loss: is bert4rec really better than sasrec?
Klenitskiy, A. and Vasilev, A · 2023
Later among the works it cites.
Text is all you need: Learning language representations for sequential recommendation
Li, J., Wang, M., Li, J., Fu, J., Shen, X., Shang, J., and McAuley, J · 2023
Later among the works it cites.
Scaling law for recommendation models: towards general-purpose user representations
Shin, K., Kwak, H., Kim, S. Y., Ramström, M. N., Jeong, J., Ha, J.-W., and Kim, K.-M · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding, 2023
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2023
Later among the works it cites.
Transact: Transformer-based realtime user action model for recommendation at pinterest
Xia, X., Eksombatchai, P., Pancha, N., Badani, D. D., Wang, P.-W., Gu, N., Joshi, S. V., Farahpour, N., Zhang, Z., and Zhai, A · 2023
Later among the works it cites.
Bytetransformer: A high-performance transformer boosted for variable-length inputs
Zhai, Y., Jiang, C., Wang, L., Jia, X., Zhang, S., Chen, Z., Liu, X., and Zhu, Y · 2023
Later among the works it cites.
Breaking the curse of quality saturation with user-centric ranking, 2023
Zhao, Z., Yang, Y., Wang, W., Liu, C., Shi, Y., Hu, W., Zhang, H., and Yang, S · 2023
Later among the works it cites.
Large language models are zero-shot rankers for recommender systems
Hou, Y., Zhang, J., Lin, Z., Lu, H., Xie, R., McAuley, J., and Zhao, W. X · 2024
Closest in time.
YaRN: Efficient context window extension of large language models
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2024
Closest in time.