Fetching the paper…
Reading the bibliography…
In recent years, there are great interests as well as challenges in applying reinforcement learning (RL) to recommendation systems (RS).
Learning dynamics model in reinforcement learning by incorporating the long term future
Ke, N. R.; Singh, A.; Touati, A.; Goyal, A.; Bengio, Y.; Parikh, D.; and Batra, D. 2019 · 1903
Earlier work this paper cites.
Challenges of real-world reinforcement learning
Dulac-Arnold, G.; Mankowitz, D.; and Hester, T. 2019 · 1904
Earlier work this paper cites.
Disentangling dynamics and returns: Value function decomposition with future prediction
Tang, H.; Hao, J.; Chen, G.; Chen, P.; Meng, Z.; Yang, Y.; and Wang, L. 2019 · 1905
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016 · 1937
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S. 1991 · 1991
Earlier work this paper cites.
An MDP-based recommender system
Shani, G.; Heckerman, D.; and Brafman, R. I. 2005 · 2005
Earlier work this paper cites.
Ad click prediction: a view from the trenches
McMahan, H. B.; Holt, G.; Sculley, D.; Young, M.; Ebner, D.; Grady, J.; Nie, L.; Phillips, T.; Davydov, E.; Golovin, D.; et al. 2013 · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013 · 2013
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2015 · 2015
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z.; Schaul, T.; Hessel, M.; Van Hasselt, H.; Lanctot, M.; and De Freitas, N. 2015 · 2015
Earlier work this paper cites.
Wide & deep learning for recommender systems
Cheng, H.-T.; Koc, L.; Harmsen, J.; Shaked, T.; Chandra, T.; Aradhye, H.; Anderson, G.; Corrado, G.; Chai, W.; Ispir, M.; et al. 2016 · 2016
Earlier work this paper cites.
Deep successor reinforcement learning
Kulkarni, T. D.; Saeedi, A.; Gautam, S.; and Gershman, S. J. 2016 · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M.; Wolski, F.; Ray, A.; Schneider, J.; Fong, R.; Welinder, P.; McGrew, B.; Tobin, J.; Abbeel, O. P.; and Zaremba, W. 2017 · 2017
Cited alongside, same era.
When recurrent neural networks meet the neighborhood for session-based recommendation
Jannach, D.; and Ludewig, M. 2017 · 2017
Cited alongside, same era.
What uncertainties do we need in bayesian deep learning for computer vision?
Kendall, A.; and Gal, Y. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Deep reinforcement learning for list-wise recommendations
Zhao, X.; Zhang, L.; Xia, L.; Ding, Z.; Yin, D.; and Tang, J. 2017 · 2017
Cited alongside, same era.
Reinforcement learning to rank in e-commerce search engine: Formalization, analysis, and application
Hu, Y.; Da, Q.; Zeng, A.; Yu, Y.; and Xu, Y. 2018 · 2018
Later among the works it cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Kendall, A.; Gal, Y.; and Cipolla, R. 2018 · 2018
Later among the works it cites.
Deep dyna-q: Integrating planning for task-completion dialogue policy learning
Peng, B.; Li, X.; Gao, J.; Liu, J.; Wong, K.-F.; and Su, S.-Y. 2018 · 2018
Later among the works it cites.
Deep reinforcement learning for page-wise recommendations
Zhao, X.; Xia, L.; Zhang, L.; Ding, Z.; Yin, D.; and Tang, J. 2018 · 2018
Later among the works it cites.
DRN: A deep reinforcement learning framework for news recommendation
Zheng, G.; Zhang, F.; Zheng, Z.; Xiang, Y.; Yuan, N. J.; Xie, X.; and Li, Z. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Colas, C.; Fournier, P.; Sigaud, O.; Chetouani, M.; and Oudeyer, P.-Y. 2018 · 2018
Cited alongside, same era.
Model-based value estimation for efficient model-free reinforcement learning
Feinberg, V.; Wan, A.; Stoica, I.; Jordan, M. I.; Gonzalez, J. E.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S.; Van Hoof, H.; and Meger, D. 2018 · 2018
Cited alongside, same era.
Ha, D.; and Schmidhuber, J. 2018 · 2018
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Hafner, D.; Lillicrap, T.; Fischer, I.; Villegas, R.; Ha, D.; Lee, H.; and Davidson, J. 2018 · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Henderson, P.; Islam, R.; Bachman, P.; Pineau, J.; Precup, D.; and Meger, D. 2018 · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M.; Modayil, J.; Van Hasselt, H.; Schaul, T.; Ostrovski, G.; Dabney, W.; Horgan, D.; Piot, B.; Azar, M.; and Silver, D. 2018 · 2018
Cited alongside, same era.
A Model-Based Reinforcement Learning with Adversarial Training for Online Recommendation
Bai, X.; Guan, J.; and Wang, H. 2019 · 2019
Later among the works it cites.
SlateQ: A tractable decomposition for reinforcement learning with recommendation sets
Ie, E.; Jain, V.; Wang, J.; Narvekar, S.; Agarwal, R.; Wu, R.; Cheng, H.-T.; Chandra, T.; and Boutilier, C. 2019 · 2019
Later among the works it cites.
Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning
Shi, J.-C.; Yu, Y.; Da, Q.; Chen, S.-Y.; and Zeng, A.-X. 2019 · 2019
Later among the works it cites.
Deep interest evolution network for click-through rate prediction
Zhou, G.; Mou, N.; Fan, Y.; Pi, Q.; Bian, W.; Zhou, C.; Zhu, X.; and Gai, K. 2019 · 2019
Later among the works it cites.
Reinforcement Learning to Optimize Long-term User Engagement in Recommender Systems
Zou, L.; Xia, L.; Ding, Z.; Song, J.; Liu, W.; and Yin, D. 2019 · 2019
Later among the works it cites.
Validation Set Evaluation can be Wrong: An Evaluator-Generator Approach for Maximizing Online Performance of Ranking in E-commerce
Huzhang, G.; Pang, Z.-J.; Gao, Y.; Zhou, W.-J.; Da, Q.; Zeng, A.-x.; and Yu, Y. 2020 · 2020
Later among the works it cites.
Pseudo Dyna-Q: A Reinforcement Learning Framework for Interactive Recommendation
Zou, L.; Xia, L.; Du, P.; Zhang, Z.; Bai, T.; Liu, W.; Nie, J.-Y.; and Yin, D. 2020 · 2020
Later among the works it cites.