Fetching the paper…
Reading the bibliography…
There are great interests as well as many challenges in applying reinforcement learning (RL) to recommendation systems.
Conditional logit analysis of qualitative choice behaviour
McFadden, D · 1973
Earlier work this paper cites.
Maximum score estimation of the stochastic utility model of choice
Manski, C. F · 1975
Earlier work this paper cites.
Learning from Delayed Rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Reward functions for accelerated learning
Mataric, M. J · 1994
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. J · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Reward function and initial values: Better choices for accelerated goal-directed reinforcement learning
Matignon, L., Laurent, G. J., and Fort-Piat, N. L · 2006
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E · 2010
Earlier work this paper cites.
Collaborative competitive filtering: learning recommender using context of user choice
Yang, S.-H., Long, B., Smola, A. J., Zha, H., and Zheng, Z · 2011
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Cited alongside, same era.
Gaussian processes for data-efficient learning in robotics and control
Deisenroth, M. P., Fox, D., and Rasmussen, C. E · 2015
Cited alongside, same era.
Session-based recommendations with recurrent neural networks
Hidasi, B., Karatzoglou, A., Baltrunas, L., and Tikk, D · 2015
Cited alongside, same era.
Xgboost: A scalable tree boosting system
Chen, T. and Guestrin, C · 2016
Cited alongside, same era.
Wide & deep learning for recommender systems
Cheng, H.-T., Koc, L., Harmsen, J., Shaked, T., Chandra, T., Aradhye, H., Anderson, G., Corrado, G., Chai, W., Ispir, M., Anil, R., Haque, Z., Hong, L., Jain, V., Liu, X., and Shah, H · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Deepfm: a factorization-machine based neural network for ctr prediction
Guo, H., Tang, R., Ye, Y., Li, Z., and He, X · 2017
Later among the works it cites.
When recurrent neural networks meet the neighborhood for session-based recommendation
Jannach, D. and Ludewig, M · 2017
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning with a combinatorial action space for predicting popular reddit threads
He, J., Ostendorf, M., He, X., Chen, J., Gao, J., Li, L., and Deng, L · 2016
Cited alongside, same era.
Session-based recommendations with recurrent neural networks
Hidasi, B., Karatzoglou, A., Baltrunas, L., and Tikk, D · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Cited alongside, same era.
Model-free imitation learning with policy optimization
Ho, J., Gupta, J. K., and Ermon, S · 2016
Cited alongside, same era.
Deep reinforcement learning for page-wise recommendations
Zhao, X., Xia, L., Zhang, L., Ding, Z., Yin, D., and Tang, J
Cited in the paper.
Deep reinforcement learning for list-wise recommendations
Zhao, X., Zhang, L., Ding, Z., Yin, D., Zhao, Y., and Tang, J
Cited in the paper.
Clavera, I., Nagabandi, A., Fearing, R. S., Abbeel, P., Levine, S., and Finn, C · 2018
Closest in time.
Behavioral cloning from observation
Torabi, F., Warnell, G., and Stone, P · 2018
Closest in time.
Drn: A deep reinforcement learning framework for news recommendation
Zheng, G., Zhang, F., Zheng, Z., Xiang, Y., Yuan, N. J., Xie, X., and Li, Z · 2018
Closest in time.
Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning
Shi, J.-C., Yu, Y., Da, Q., Chen, S.-Y., and Zeng, A.-X · 2019
Closest in time.