Fetching the paper…
Reading the bibliography…
Most of the existing deep reinforcement learning (RL) approaches for session-based recommendations either rely on costly online interactions with real users, or rely on potentially biased rule-based or data-driven user-behavior models for learning.
E. Ie, V. Jain, J. Wang, S. Navrekar, R. Agarwal, R. Wu, H.-T. Cheng, M. Lustman, V. Gatto, P. Covington, et al · 1905
Earlier work this paper cites.
Recsim: A configurable simulation platform for recommender systems
E. Ie, C.-w. Hsu, M. Mladenov, V. Jain, S. Narvekar, J. Wang, R. Wu, and C. Boutilier · 1909
Earlier work this paper cites.
Benchmarking batch deep reinforcement learning algorithms
S. Fujimoto, E. Conti, M. Ghavamzadeh, and J. Pineau · 1910
Earlier work this paper cites.
Robust estimation of a location parameter
P. J. Huber et al · 1964
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
L.-J. Lin · 1992
Earlier work this paper cites.
An mdp-based recommender system
G. Shani, D. Heckerman, and R. I. Brafman · 2005
Earlier work this paper cites.
Z. Wang, A. Novikov, K. Żołna, J. T. Springenberg, S. Reed, B. Shahriari, N. Siegel, J. Merel, C. Gulcehre, N. Heess, et al · 2006
Earlier work this paper cites.
Nonparametric return distribution approximation for reinforcement learning
T. Morimura, M. Sugiyama, H. Kashima, H. Hachiya, and T. Tanaka · 2010
Earlier work this paper cites.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
L. Li, W. Chu, J. Langford, and X. Wang · 2011
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2011
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder–decoder approaches
K. Cho, B. van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
word2vec explained: deriving mikolov et al.’s negative-sampling word-embedding method
Y. Goldberg and O. Levy · 2014
Earlier work this paper cites.
Factored mdps for detecting topics of user sessions
M. Tavakol and U. Brefeld · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
R. S. Sutton, A. R. Mahmood, and M. White · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Cited alongside, same era.
When recurrent neural networks meet the neighborhood for session-based recommendation
D. Jannach and M. Ludewig · 2017
Cited alongside, same era.
Off-policy evaluation for slate recommendation
A. Swaminathan, A. Krishnamurthy, A. Agarwal, M. Dudik, J. Langford, D. Jose, and I. Zitouni · 2017
Generative adversarial user model for reinforcement learning based recommendation system
X. Chen, S. Li, H. Li, S. Jiang, Y. Qi, and L. Song · 2019
Later among the works it cites.
Sequence and time aware neighborhood for session-based recommendations: Stan
D. Garg, P. Gupta, P. Malhotra, L. Vig, and G. Shroff · 2019
Later among the works it cites.
Niser: Normalized item and session representations with graph neural networks
P. Gupta, D. Garg, P. Malhotra, L. Vig, and G. Shroff · 2019
Later among the works it cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep reinforcement learning for list-wise recommendations
X. Zhao, L. Zhang, L. Xia, Z. Ding, D. Yin, and J. Tang · 2017
Cited alongside, same era.
Recurrent neural networks with top-k gains for session-based recommendations
B. Hidasi and A. Karatzoglou · 2018
Cited alongside, same era.
Stamp: short-term attention/memory priority model for session-based recommendation
Q. Liu, Y. Zeng, R. Mokhosi, and H. Zhang · 2018
Cited alongside, same era.
Evaluation of session-based recommendation algorithms
M. Ludewig and D. Jannach · 2018
Cited alongside, same era.
Unbiased offline recommender evaluation for missing-not-at-random implicit feedback
L. Yang, Y. Cui, Y. Xuan, C. Wang, S. Belongie, and D. Estrin · 2018
Cited alongside, same era.
Drn: A deep reinforcement learning framework for news recommendation
G. Zheng, F. Zhang, Z. Zheng, Y. Xiang, N. J. Yuan, X. Xie, and Z. Li · 2018
Cited alongside, same era.
Later among the works it cites.
Session-based recommendation with graph neural networks
S. Wu, Y. Tang, Y. Zhu, L. Wang, X. Xie, and T. Tan · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
R. Agarwal, D. Schuurmans, and M. Norouzi · 2020
Closest in time.
Causality and batch reinforcement learning: Complementary approaches to planning in unknown domains
J. Bannon, B. Windsor, W. Song, and T. Li · 2020
Closest in time.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Closest in time.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Closest in time.
Hyperparameter selection for offline reinforcement learning
T. L. Paine, C. Paduraru, A. Michi, C. Gulcehre, K. Zolna, A. Novikov, Z. Wang, and N. de Freitas · 2020
Closest in time.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2062
Closest in time.