Fetching the paper…
Reading the bibliography…
We study session-based recommendation scenarios where we want to recommend items to users during sequential interactions to improve their long-term utility.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Nachum, O.; Chow, Y.; Dai, B.; and Li, L. 2019 · 1906
Earlier work this paper cites.
Composite Q-learning: Multi-scale Q-function Decomposition and Separable Optimization
Kalweit, G.; Huegle, M.; and Boedecker, J. 2019 · 1909
Earlier work this paper cites.
Returning is believing: Optimizing long-term user engagement in recommender systems
Wu, Q.; Wang, H.; Hong, L.; and Shi, Y. 2017 · 1936
Earlier work this paper cites.
Discrete dynamic programming
Blackwell, D. 1962 · 1962
Earlier work this paper cites.
Experiments in nonconvex optimization: stochastic approximation with function smoothing and simulated annealing
Styblinski, M.; and Tang, T.-S. 1990 · 1990
Earlier work this paper cites.
Neuro-dynamic programming: an overview
Bertsekas, D. P.; and Tsitsiklis, J. N. 1995 · 1995
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D. 2000 · 2000
Earlier work this paper cites.
Blackwell optimality
Hordijk, A.; and Yushkevich, A. A. 2002 · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S.; and Langford, J. 2002 · 2002
Earlier work this paper cites.
Adaptive Estimator Selection for Off-Policy Evaluation
Su, Y.; Srinath, P.; and Krishnamurthy, A. 2020 · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M. 2003 · 2003
Earlier work this paper cites.
An MDP-based recommender system
Shani, G.; Heckerman, D.; Brafman, R. I.; and Boutilier, C. 2005 · 2005
Earlier work this paper cites.
Clinical data based optimal STI strategies for HIV: a reinforcement learning approach
Ernst, D.; Stan, G.-B.; Goncalves, J.; and Wehenkel, L. 2006 · 2006
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Kumar, A.; Zhou, A.; Tucker, G.; and Levine, S. 2020 · 2006
Cited alongside, same era.
Provably good batch reinforcement learning without great exploration
Liu, Y.; Swaminathan, A.; Agarwal, A.; and Brunskill, E. 2020 · 2007
Cited alongside, same era.
Truncated importance sampling
Ionides, E. L. 2008 · 2008
Cited alongside, same era.
Batch-Constrained Distributional Reinforcement Learning for Session-based Recommendation
Garg, D.; Gupta, P.; Malhotra, P.; Vig, L.; and Shroff, G. 2020 · 2012
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C.; and Brunskill, E. 2015 · 2015
DRN: A deep reinforcement learning framework for news recommendation
Zheng, G.; Zhang, F.; Zheng, Z.; Xiang, Y.; Yuan, N. J.; Xie, X.; and Li, Z. 2018 · 2018
Later among the works it cites.
Safe policy improvement with baseline bootstrapping
Laroche, R.; Trichelair, P.; and Des Combes, R. T. 2019 · 2019
Later among the works it cites.
Policy gradients for contextual recommendations
Pan, F.; Cai, Q.; Tang, P.; Zhuang, F.; and He, Q. 2019 · 2019
Later among the works it cites.
Separating value functions across time-scales
Romoff, J.; Henderson, P.; Touati, A.; Brunskill, E.; Pineau, J.; and Ollivier, Y. 2019 · 2019
Later among the works it cites.
Policy Improvement via Imitation of Multiple Oracles
Cheng, C.-A.; Kolobov, A.; and Agarwal, A. 2020 · 2020
Later among the works it cites.
Personalized heartsteps: A reinforcement learning algorithm for optimizing physical activity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The dependence of effective planning horizon on model accuracy
Jiang, N.; Kulesza, A.; Singh, S.; and Lewis, R. 2015 · 2015
Cited alongside, same era.
When recurrent neural networks meet the neighborhood for session-based recommendation
Jannach, D.; and Ludewig, M. 2017 · 2017
Cited alongside, same era.
Neural attentive session-based recommendation
Li, J.; Ren, P.; Chen, Z.; Ren, Z.; Lian, T.; and Ma, J. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Evaluation of session-based recommendation algorithms
Ludewig, M.; and Jannach, D. 2018 · 2018
Cited alongside, same era.
Rohde, D.; Bonner, S.; Dunlop, T.; Vasile, F.; and Karatzoglou, A. 2018 · 2018
Cited alongside, same era.
Truncated horizon policy search: Combining reinforcement learning & imitation learning
Sun, W.; Bagnell, J. A.; and Boots, B. 2018 · 2018
Cited alongside, same era.
Liao, P.; Greenewald, K.; Klasnja, P.; and Murphy, S. 2020 · 2020
Later among the works it cites.
Off-policy learning in two-stage recommender systems
Ma, J.; Zhao, Z.; Yi, X.; Yang, J.; Chen, M.; Tang, J.; Hong, L.; and Chi, E. H. 2020a · 2020
Later among the works it cites.
Non-Stationary Delayed Bandits with Intermediate Observations
Vernade, C.; Gyorgy, A.; and Mann, T. 2020 · 2020
Later among the works it cites.
Preference-aware mask for session-based recommendation with bidirectional transformer
Zhang, Y.; Zhao, P.; Guan, Y.; Chen, L.; Bian, K.; Song, L.; Cui, B.; and Li, X. 2020 · 2020
Later among the works it cites.
Heuristic-Guided Reinforcement Learning
Cheng, C.; Kolobov, A.; and Swaminathan, A. 2021 · 2021
Closest in time.
Near-Optimal Offline Reinforcement Learning via Double Variance Reduction
Yin, M.; Bai, Y.; and Wang, Y.-X. 2021 · 2021
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S.; Meger, D.; and Precup, D. 2019 · 2062
Closest in time.