Fetching the paper…
Reading the bibliography…
The wide popularity of short videos on social media poses new opportunities and challenges to optimize recommender systems on the video-sharing platforms.
Asynchronous methods for deep reinforcement learning. In International conference on machine learning . PMLR, 1928–1937
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 1930–1939
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi. 2018 · 1939
Earlier work this paper cites.
Solution procedures for vector criterion Markov decision processes
C Ch White, CC III WHITE, and KIM KW. 1980 · 1980
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Multitask learning
Rich Caruana. 1997 · 1997
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup. 2000 · 2000
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation. In ICML . 417–424
Doina Precup, Richard S Sutton, and Sanjoy Dasgupta. 2001 · 2001
Earlier work this paper cites.
Lyapunov design for safe reinforcement learning
Theodore J Perkins and Andrew G Barto. 2002 · 2002
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández. 2015 · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015 · 2015
Earlier work this paper cites.
The self-normalized estimator for counterfactual learning
Adith Swaminathan and Thorsten Joachims. 2015 · 2015
Earlier work this paper cites.
Joint multi-grain topic sentiment: modeling semantic aspects for online reviews
Md Hijbul Alam, Woo-Jong Ryu, and SangKeun Lee. 2016 · 2016
Earlier work this paper cites.
Wide & deep learning for recommender systems. In Proceedings of the 1st workshop on deep learning for recommender systems . 7–10
Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, et al · 2016
Earlier work this paper cites.
Multi-objective deep reinforcement learning
Hossam Mossalam, Yannis M Assael, Diederik M Roijers, and Shimon Whiteson. 2016 · 2016
Earlier work this paper cites.
Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach. In 2016 38th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC) . IEEE, 2978–2981
Shamim Nemati, Mohammad M Ghassemi, and Gari D Clifford. 2016 · 2016
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone. 2017 · 2017
Cited alongside, same era.
Reverse curriculum generation for reinforcement learning. In Conference on robot learning . PMLR, 482–495
Carlos Florensa, David Held, Markus Wulfmeier, Michael Zhang, and Pieter Abbeel. 2017 · 2017
Cited alongside, same era.
DeepFM: a factorization-machine based neural network for CTR prediction
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017 · 2017
Cited alongside, same era.
Multi-task deep reinforcement learning with popart. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 3796–3803
Matteo Hessel, Hubert Soyer, Lasse Espeholt, Wojciech Czarnecki, Simon Schmitt, and Hado van Hasselt. 2019 · 2019
Later among the works it cites.
A deep reinforcement learning approach to proactive content pushing and recommendation for mobile users
Dong Liu and Chenyang Yang. 2019 · 2019
Later among the works it cites.
Reinforcement knowledge graph reasoning for explainable recommendation. In Proceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval . 285–294
Yikun Xian, Zuohui Fu, Shan Muthukrishnan, Gerard De Melo, and Yongfeng Zhang. 2019 · 2019
Later among the works it cites.
Reinforcement learning to optimize long-term user engagement in recommender systems. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 2810–2818
Lixin Zou, Long Xia, Zhuoye Ding, Jiaxing Song, Weidong Liu, and Dawei Yin. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiangyu Zhao, Liang Zhang, Long Xia, Zhuoye Ding, Dawei Yin, and Jiliang Tang. 2017 · 2017
Cited alongside, same era.
Stabilizing reinforcement learning in dynamic environment with application to online recommendation. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 1187–1196
Shi-Yong Chen, Yang Yu, Qing Da, Jun Tan, Hai-Kuan Huang, and Hai-Hong Tang. 2018 · 2018
Cited alongside, same era.
Safe exploration in continuous action spaces
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa. 2018 · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods. In International conference on machine learning . PMLR, 1587–1596
Scott Fujimoto, Herke Hoof, and David Meger. 2018 · 2018
Cited alongside, same era.
Learning by playing solving sparse reward tasks from scratch. In International Conference on Machine Learning . PMLR, 4344–4353
Martin Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Wiele, Vlad Mnih, Nicolas Heess, and Jost Tobias Springenberg. 2018 · 2018
Cited alongside, same era.
Multi-task learning as multi-objective optimization
Ozan Sener and Vladlen Koltun. 2018 · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Cited alongside, same era.
Modular multi-objective deep reinforcement learning with decision values. In 2018 Federated conference on computer science and information systems (FedCSIS) . IEEE, 85–93
Tomasz Tajmajer. 2018 · 2018
Cited alongside, same era.
Off-policy learning in two-stage recommender systems. In Proceedings of The Web Conference 2020 . 463–473
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang, Minmin Chen, Jiaxi Tang, Lichan Hong, and Ed H Chi. 2020 · 2020
Later among the works it cites.
AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine. 2020 · 2020
Later among the works it cites.
A multi-objective deep reinforcement learning framework
Thanh Thi Nguyen, Ngoc Duy Nguyen, Peter Vamplew, Saeid Nahavandi, Richard Dazeley, and Chee Peng Lim. 2020 · 2020
Later among the works it cites.
Search-based user interest modeling with lifelong sequential behavior data for click-through rate prediction. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management . 2685–2692
Qi Pi, Guorui Zhou, Yujing Zhang, Zhe Wang, Lejian Ren, Ying Fan, Xiaoqiang Zhu, and Kun Gai. 2020 · 2020
Later among the works it cites.
Reinforcement learning based recommender systems: A survey
M Mehdi Afsar, Trafford Crump, and Behrouz Far. 2021 · 2021
Later among the works it cites.
Reinforcement Recommendation with User Multi-aspect Preference. In Proceedings of the Web Conference 2021 . 425–435
Xu Chen, Yali Du, Long Xia, and Jun Wang. 2021 · 2021
Later among the works it cites.
Towards Long-term Fairness in Recommendation. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining . 445–453
Yingqiang Ge, Shuchang Liu, Ruoyuan Gao, Yikun Xian, Yunqi Li, Xiangyu Zhao, Changhua Pei, Fei Sun, Junfeng Ge, Wenwu Ou, et al · 2021
Later among the works it cites.
Policy learning with constraints in model-free reinforcement learning: A survey. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence
Yongshuai Liu, Avishai Halev, and Xin Liu. 2021 · 2021
Later among the works it cites.
Dusan Stamenkovic, Alexandros Karatzoglou, Ioannis Arapakis, Xin Xin, and Kleomenis Katevas. 2021 · 2021
Later among the works it cites.
Kuaishou Reports $3.2 Billion Total Revenue, Expanded Net Loss in Q3
[n. d.] · 2022
Closest in time.
TikTok User Statistics (2022)
[n. d.] · 2022
Closest in time.