Fetching the paper…
Reading the bibliography…
The click-through rate (CTR) prediction task is to predict whether a user will click on the recommended item.
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
You, Y.; Li, J.; Reddi, S. J.; Hseu, J.; Kumar, S.; Bhojanapalli, S.; Song, X.; Demmel, J.; Keutzer, K.; and Hsieh, C.-J. 2020 · 1904
Earlier work this paper cites.
Why Gradient Clipping Accelerates Training: A Theoretical Justification for Adaptivity
Zhang, J.; He, T.; Sra, S.; and Jadbabaie, A. 2020 · 1905
Earlier work this paper cites.
Deep Learning Recommendation Model for Personalization and Recommendation Systems
Naumov, M.; Mudigere, D.; Shi, H.-J. M.; Huang, J.; Sundaraman, N.; Park, J.; Wang, X.; Gupta, U.; Wu, C.-J.; Azzolini, A. G.; Dzhulgakov, D.; Mallevich, A.; Cherniavskii, I.; Lu, Y.; Krishnamoorthi, R.; Yu, A.; Kondratenko, V.; Pereira, S.; Chen, X.; Chen, W.; Rao, V.; Jia, B.; Xiong, L.; and Smelyanskiy, M. 2019 · 1906
Earlier work this paper cites.
Scale MLPerf-0.6 models on Google TPU-v3 Pods
Kumar, S.; Bitorff, V.; Chen, D.; Chou, C.-H.; Hechtman, B. A.; Lee, H.; Kumar, N.; Mattson, P.; Wang, S.; Wang, T.; Xu, Y.; and Zhou, Z. 2019 · 1909
Earlier work this paper cites.
Mattson, P.; Cheng, C.; Coleman, C. A.; Diamos, G. F.; Micikevicius, P.; Patterson, D.; Tang, H.; Wei, G.-Y.; Bailis, P.; Bittorf, V.; Brooks, D. M.; Chen, D.; Dutta, D.; Gupta, U.; Hazelwood, K. M.; Hock, A.; Huang, X.; Jia, B.; Kang, D.; Kanter, D.; Kumar, N.; Liao, J.; Ma, G.; Narayanan, D.; Oguntebi, T.; Pekhimenko, G.; Pentecost, L.; Reddi, V. J.; Robie, T.; John, T. S.; Wu, C.-J.; Xu, L.; Young, C.; and Zaharia, M. A. 2020 · 1910
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N.; Hinton, G. E.; Krizhevsky, A.; Sutskever, I.; and Salakhutdinov, R. 2014 · 1958
Earlier work this paper cites.
Wang, Z.; She, Q.; Zhang, P.; and Zhang, J. 2020 · 2006
Earlier work this paper cites.
User experience of mobile internet: analysis and recommendations
Kaasinen, E.; Roto, V.; Roloff, K.; Väänänen-Vainio-Mattila, K.; Vainio, T.; Maehr, W.; Joshi, D.; and Shrestha, S. 2009 · 2009
Earlier work this paper cites.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
Duchi, J. C.; Hazan, E.; and Singer, Y. 2010 · 2010
Earlier work this paper cites.
Factorization Machines
Rendle, S. 2010 · 2010
Earlier work this paper cites.
Ad click prediction: a view from the trenches
McMahan, H. B.; Holt, G.; Sculley, D.; Young, M.; Ebner, D.; Grady, J.; Nie, L.; Phillips, T.; Davydov, E.; Golovin, D.; Chikkerur, S.; Liu, D.; Wattenberg, M.; Hrafnkelsson, A. M.; Boulos, T.; and Kubica, J. 2013 · 2013
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A. 2014 · 2014
Earlier work this paper cites.
Display Advertising Challenge
Labs, C. 2014 · 2014
Earlier work this paper cites.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems
Abadi, M.; Agarwal, A.; Barham, P.; Brevdo, E.; Chen, Z.; Citro, C.; Corrado, G. S.; Davis, A.; Dean, J.; Devin, M.; Ghemawat, S.; Goodfellow, I.; Harp, A.; Irving, G.; Isard, M.; Jia, Y.; Jozefowicz, R.; Kaiser, L.; Kudlur, M.; Levenberg, J.; Mané, D.; Monga, R.; Moore, S.; Murray, D.; Olah, C.; Schuster, M.; Shlens, J.; Steiner, B.; Sutskever, I.; Talwar, K.; Tucker, P.; Vanhoucke, V.; Vasudevan, V.; Viégas, F.; Vinyals, O.; Warden, P.; Wattenberg, M.; Wicke, M.; Yu, Y.; and Zheng, X. 2015 · 2015
Earlier work this paper cites.
Avazu Click-Through Rate Prediction
Avazu. 2015 · 2015
Earlier work this paper cites.
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Wide & Deep Learning for Recommender Systems
Cheng, H.-T.; Koc, L.; Harmsen, J.; Shaked, T.; Chandra, T.; Aradhye, H. B.; Anderson, G.; Corrado, G. S.; Chai, W.; Ispir, M.; Anil, R.; Haque, Z.; Hong, L.; Jain, V.; Liu, X.; and Shah, H. 2016 · 2016
Earlier work this paper cites.
Deep Neural Networks for YouTube Recommendations
Covington, P.; Adams, J. K.; and Sargin, E. 2016 · 2016
Earlier work this paper cites.
The Netflix Recommender System: Algorithms, Business Value, and Innovation
Gomez-Uribe, C.; and Hunt, N. 2016 · 2016
Earlier work this paper cites.
DiFacto: Distributed Factorization Machines
Li, M.; Liu, Z.; Smola, A.; and Wang, Y.-X. 2016 · 2016
Cited alongside, same era.
Product-Based Neural Networks for User Response Prediction
Qu, Y.; Cai, H.; Ren, K.; Zhang, W.; Yu, Y.; Wen, Y.; and Wang, J. 2016 · 2016
Cited alongside, same era.
Deep Learning over Multi-field Categorical Data – A Case Study on User Response Prediction
Zhang, W.; Du, T.; and Wang, J. 2016 · 2016
Cited alongside, same era.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Goyal, P.; Dollár, P.; Girshick, R. B.; Noordhuis, P.; Wesolowski, L.; Kyrola, A.; Tulloch, A.; Jia, Y.; and He, K. 2017 · 2017
Cited alongside, same era.
Hoffer, E.; Hubara, I.; and Soudry, D. 2017 · 2017
Cited alongside, same era.
AIBox: CTR Prediction Model Training on a Single Node
Zhao, W.; Zhang, J.; Xie, D.; Qian, Y.; Jia, R.; and Li, P. 2019 · 2019
Later among the works it cites.
Deep Interest Evolution Network for Click-Through Rate Prediction
Zhou, G.; Mou, N.; Fan, Y.; Pi, Q.; Bian, W.; Zhou, C.; Zhu, X.; and Gai, K. 2019 · 2019
Later among the works it cites.
Temporal-Contextual Recommendation in Real-Time
Ma, Y.; Narayanaswamy, B.; Lin, H.; and Ding, H. 2020 · 2020
Later among the works it cites.
A Survey of Online Advertising Click-Through Rate Prediction Models
Wang, X. 2020 · 2020
Later among the works it cites.
Kraken: memory-efficient continual learning for large-scale real-time recommendations
Xie, M.; Ren, K.; Lu, Y.; Yang, G.; Xu, Q.; Wu, B.; Lin, J.; Ao, H.; Xu, W.; and Shu, J. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
DeepCTR: Easy-to-use,Modular and Extendible package of deep-learning based CTR models
Shen, W. 2017 · 2017
Cited alongside, same era.
Deep & Cross Network for Ad Click Predictions
Wang, R.; Fu, B.; Fu, G.; and Wang, M. 2017 · 2017
Cited alongside, same era.
Large Batch Training of Convolutional Networks
You, Y.; Gitman, I.; and Ginsburg, B. 2017 · 2017
Cited alongside, same era.
Deep learning using rectified linear units (relu)
Agarap, A. F. 2018 · 2018
Cited alongside, same era.
Evolution of the GPU Device widely used in AI and Massive Parallel Processing
Baji, T. 2018 · 2018
Cited alongside, same era.
DeepFM: An End-to-End Wide & Deep Learning Framework for CTR Prediction
Guo, H.; Tang, R.; Ye, Y.; Li, Z.; He, X.; and Dong, Z. 2018 · 2018
Cited alongside, same era.
xDeepFM: Combining Explicit and Implicit Feature Interactions for Recommender Systems
Lian, J.; Zhou, X.; Zhang, F.; Chen, Z.; Xie, X.; and zhong Sun, G. 2018 · 2018
Cited alongside, same era.
Adnan, M. 2021 · 2021
Later among the works it cites.
Accelerating Recommendation System Training by Leveraging Popular Choices
Adnan, M.; Maboud, Y. E.; Mahajan, D.; and Nair, P. J. 2021 · 2021
Later among the works it cites.
High-performance large-scale image recognition without normalization
Brock, A.; De, S.; Smith, S. L.; and Simonyan, K. 2021 · 2021
Later among the works it cites.
Improved CTR Prediction Algorithm based on LSTM and Attention
Chen, Q.; and Li, D. 2021 · 2021
Later among the works it cites.
DeepLight: Deep Lightweight Feature Interactions for Accelerating CTR Predictions in Ad Serving
Deng, W.; Pan, J.; Zhou, T.; Flores, A.; and Lin, G. 2021 · 2021
Later among the works it cites.
Mixed Dimension Embeddings with Application to Memory-Efficient Recommendation Systems
Ginart, A. A.; Naumov, M.; Mudigere, D.; Yang, J.; and Zou, J. Y. 2021 · 2021
Later among the works it cites.
Large-Scale Deep Learning Optimizations: A Comprehensive Survey
He, X.; Xue, F.; Ren, X.; and You, Y. 2021 · 2021
Later among the works it cites.
Memorize, Factorize, or be Naïve: Learning Optimal Feature Interaction Methods for CTR Prediction
Lyu, F.; Tang, X.; Guo, H.; Tang, R.; He, X.; Zhang, R.; and Liu, X. 2021 · 2021
Later among the works it cites.
Stability and convergence of stochastic gradient clipping: Beyond lipschitz continuity and smoothness
Mai, V. V.; and Johansson, M. 2021 · 2021
Later among the works it cites.
HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework
Miao, X.; Zhang, H.; Shi, Y.; Nie, X.; Yang, Z.; Tao, Y.; and Cui, B. 2021 · 2021
Later among the works it cites.
Software-Hardware Co-design for Fast and Scalable Training of Deep Learning Recommendation Models
Mudigere, D.; Hao, Y.; Huang, J.; Jia, Z.; Tulloch, A.; Sridharan, S.; Liu, X.; Ozdal, M.; Nie, J.; Park, J.; Luo, L.; Yang, J. A.; Gao, L.; Ivchenko, D.; Basant, A.; Hu, Y.; Yang, J.; Ardestani, E. K.; Wang, X.; Komuravelli, R.; Chu, C.-H.; Yilmaz, S.; Li, H.; Qian, J.; Feng, Z.; Ma, Y.-A.; Yang, J.; Wen, E.; Li, H.; Yang, L.; Sun, C.; Zhao, W.; Melts, D.; Dhulipala, K.; Kishore, K. G.; Graf, T.; Eisenman, A.; Matam, K. K.; Gangidi, A.; Chen, G. J.; Krishnan, M.; Nayak, A.; Nair, K.; Muthiah, B.; khorashadi, M.; Bhattacharya, P.; Lapukhov, P.; Naumov, M.; Mathews, A. S.; Qiao, L.; Smelyanskiy, M.; Jia, B.; and Rao, V. 2021 · 2021
Later among the works it cites.
DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems
Wang, R.; Shivanna, R.; Cheng, D. Z.; Jain, S.; Lin, D.; Hong, L.; and Chi, E. H. 2021 · 2021
Later among the works it cites.
MaskNet: Introducing Feature-Wise Multiplication to CTR Ranking Models by Instance-Guided Mask
Wang, Z.; She, Q.; and Zhang, J. 2021 · 2021
Later among the works it cites.
Open Benchmarking for Click-Through Rate Prediction
Zhu, J.; Liu, J.; Yang, S.; Zhang, Q.; and He, X. 2021 · 2021
Later among the works it cites.