Fetching the paper…
Reading the bibliography…
How might we design Reinforcement Learning (RL)-based recommenders that encourage aligning user trajectories with the underlying user satisfaction? Three research questions are key: (1) measuring user satisfaction, (2) combatting sparsity of satisfaction signals, and (3) adapting the training of the recommender agent to maximize satisfaction.
Slateq: A tractable decomposition for reinforcement learning with recommendation sets
Ie, E., Jain, V., Wang, J., Narvekar, S., Agarwal, R., Wu, R., Cheng, H.-T., Chandra, T., and Boutilier, C · 1905
Earlier work this paper cites.
Measurement and control of response bias
Paulhus, D. L · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Robot shaping: Developing autonomous agents through learning
Dorigo, M. and Colombetti, M · 1994
Earlier work this paper cites.
Reward functions for accelerated learning
Mataric, M. J · 1994
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Schaal, S · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
The foundations of cost-sensitive learning
Elkan, C · 2001
Earlier work this paper cites.
A survey on transfer learning
Pan, S. J. and Yang, Q · 2009
Earlier work this paper cites.
Beyond dwell time: estimating document relevance from cursor movements and other post-click searcher behavior
Guo, Q. and Agichtein, E · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Simple and scalable response prediction for display advertising
Chapelle, O., Manavoglu, E., and Rosales, R · 2014
Cited alongside, same era.
Beyond clicks: dwell time for personalization
Yi, X., Hong, L., Zhong, E., Liu, N. N., and Rajan, S · 2014
Cited alongside, same era.
Deep reinforcement learning in large discrete action spaces
Dulac-Arnold, G., Evans, R., van Hasselt, H., Sunehag, P., Lillicrap, T., Hunt, J., Mann, T., Weber, T., Degris, T., and Coppin, B · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Cited alongside, same era.
Deep reinforcement learning based recommendation with explicit user-item interactions modeling
Liu, F., Tang, R., Li, X., Zhang, W., Ye, Y., Chen, H., Guo, H., and Zhang, Y · 2018
Later among the works it cites.
Towards a fair marketplace: Counterfactual evaluation of the trade-off between relevance, fairness & satisfaction in recommendation systems
Mehrotra, R., McInerney, J., Bouchard, H., Lalmas, M., and Diaz, F · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Deep reinforcement learning for page-wise recommendations
Zhao, X., Xia, L., Zhang, L., Ding, Z., Yin, D., and Tang, J · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep neural networks for youtube recommendations
Covington, P., Adams, J., and Sargin, E · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Gu, S., Holly, E., Lillicrap, T., and Levine, S · 2017
Cited alongside, same era.
Inverse reward design
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J., and Dragan, A · 2017
Cited alongside, same era.
Latent cross: Making use of context in recurrent recommender systems
Beutel, A., Covington, P., Jain, S., Xu, C., Li, J., Gatto, V., and Chi, E. H · 2018
Cited alongside, same era.
Understanding and evaluating user satisfaction with music discovery
Garcia-Gathright, J., St. Thomas, B., Hosey, C., Nazari, Z., and Diaz, F · 2018
Cited alongside, same era.
Drn: A deep reinforcement learning framework for news recommendation
Zheng, G., Zhang, F., Zheng, Z., Xiang, Y., Yuan, N. J., Xie, X., and Li, Z · 2018
Later among the works it cites.
Top-k off-policy correction for a reinforce recommender system
Chen, M., Beutel, A., Covington, P., Jain, S., Belletti, F., and Chi, E. H · 2019
Later among the works it cites.
Metrics, engagement & personalization
Lalmas, M · 2019
Later among the works it cites.
Jointly leveraging intent and interaction signals to predict user satisfaction with slate recommendations
Mehrotra, R., Lalmas, M., Kenney, D., Lim-Meng, T., and Hashemian, G · 2019
Later among the works it cites.
Leveraging post-click feedback for content recommendations
Wen, H., Yang, L., and Estrin, D · 2019
Later among the works it cites.
Hierarchical reinforcement learning for course recommendation in moocs
Zhang, J., Hao, B., Chen, B., Li, C., Chen, H., and Sun, J · 2019
Later among the works it cites.