Fetching the paper…
Reading the bibliography…
Recommender systems play an important role in many content platforms.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts. In ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . 1930–1939
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi. 2018 · 1939
Earlier work this paper cites.
Multitask learning
Rich Caruana. 1997 · 1997
Earlier work this paper cites.
Large scale online learning
Léon Bottou and Yann Cun. 2003 · 2003
Earlier work this paper cites.
Convolutional deep belief networks on cifar-10
Alex Krizhevsky and Geoff Hinton. 2010 · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer. 2011 · 2011
Earlier work this paper cites.
Online learning and online convex optimization
Shai Shalev-Shwartz et al · 2012
Earlier work this paper cites.
Ad click prediction: a view from the trenches. In Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining . 1222–1230
H Brendan McMahan, Gary Holt, David Sculley, Michael Young, Dietmar Ebner, Julian Grady, Lan Nie, Todd Phillips, Eugene Davydov, Daniel Golovin, et al · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks. In International conference on machine learning . PMLR, 1310–1318
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2013 · 2013
Earlier work this paper cites.
Deep content-based music recommendation
Aaron Van den Oord, Sander Dieleman, and Benjamin Schrauwen. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning . PMLR, 448–456
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Tensorflow: a system for large-scale machine learning.. In OSDI . Savannah, GA, USA
Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Cited alongside, same era.
Latent Cross: Making Use of Context in Recurrent Recommender Systems. In International Conference on Web Search and Data Mining . ACM, 46–54
Why Gradient Clipping Accelerates Training: A Theoretical Justification for Adaptivity. In International Conference on Learning Representations
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie. 2020 · 2020
Later among the works it cites.
High-performance large-scale image recognition without normalization. In International Conference on Machine Learning . PMLR, 1059–1071
Andy Brock, Soham De, Samuel L Smith, and Karen Simonyan. 2021 · 2021
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy M Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar. 2021 · 2021
Later among the works it cites.
A Loss Curvature Perspective on Training Instability in Deep Learning
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta, Behnam Neyshabur, David Cardoze, George Dahl, Zachary Nado, and Orhan Firat. 2021b · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Beutel, Paul Covington, Sagar Jain, Can Xu, Jia Li, Vince Gatto, and Ed H Chi. 2018 · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning . PMLR, 4596–4604
Noam Shazeer and Mitchell Stern. 2018 · 2018
Cited alongside, same era.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E. 2018 · 2018
Cited alongside, same era.
Towards neural mixture recommender for long range dependent user sequences. In The World Wide Web Conference . 1782–1793
Jiaxi Tang, Francois Belletti, Sagar Jain, Minmin Chen, Alex Beutel, Can Xu, and Ed H. Chi. 2019 · 2019
Cited alongside, same era.
Sampling-bias-corrected neural modeling for large corpus item recommendations. In Proceedings of the 13th ACM Conference on Recommender Systems . 269–277
Xinyang Yi, Ji Yang, Lichan Hong, Derek Zhiyuan Cheng, Lukasz Heldt, Aditee Kumthekar, Zhe Zhao, Li Wei, and Ed Chi. 2019 · 2019
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2019 · 2019
Cited alongside, same era.
Recommending what video to watch next: a multitask ranking system. In Proceedings of the 13th ACM Conference on Recommender Systems . 43–51
Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, and Ed Chi. 2019 · 2019
Cited alongside, same era.
Large scale video representation learning via relational graph clustering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6807–6816
Hyodong Lee, Joonseok Lee, Joe Yue-Hei Ng, and Paul Natsev. 2020 · 2020
Cited alongside, same era.
David R So, Wojciech Mańke, Hanxiao Liu, Zihang Dai, Noam Shazeer, and Quoc V Le. 2021 · 2021
Later among the works it cites.
Dcn v2: Improved deep & cross network and practical lessons for web-scale learning to rank systems. In Proceedings of the Web Conference 2021 . 1785–1797
Ruoxi Wang, Rakesh Shivanna, Derek Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed Chi. 2021 · 2021
Later among the works it cites.
On the Factory Floor: ML Engineering for Industrial-Scale Ads Recommendation Models
Rohan Anil, Sandra Gadanho, Da Huang, Nijith Jacob, Zhuoshu Li, Dong Lin, Todd Phillips, Cristina Pop, Kevin Regan, Gil I Shamir, et al · 2022
Later among the works it cites.
Understanding Scaling Laws for Recommendation Models
Newsha Ardalani, Carole-Jean Wu, Zeliang Chen, Bhargav Bhushanam, and Adnan Aziz. 2022 · 2022
Later among the works it cites.
Evolved Optimizer for Vision. In First Conference on Automated Machine Learning (Late-Breaking Workshop)
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Yao Liu, Kaiyuan Wang, Cho-Jui Hsieh, Yifeng Lu, and Quoc V Le. 2022 · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Later among the works it cites.