Fetching the paper…
Reading the bibliography…
Recommender systems predict what items a user will interact with next, based on their past interactions.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S. (2019) · 1910
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O. (2019) · 1911
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D. (2000) · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020) · 2005
Earlier work this paper cites.
Collaborative filtering and the missing at random assumption
Marlin, B. M., Zemel, R. S., Roweis, S., and Slaney, M. (2007) · 2007
Earlier work this paper cites.
Collaborative filtering for implicit feedback datasets
Hu, Y., Koren, Y., and Volinsky, C. (2008) · 2008
Earlier work this paper cites.
One-class collaborative filtering
Pan, R., Zhou, Y., Cao, B., Liu, N. N., Lukose, R., Scholz, M., and Yang, Q. (2008) · 2008
Earlier work this paper cites.
Matrix factorization techniques for recommender systems
Koren, Y., Bell, R., and Volinsky, C. (2009) · 2009
Earlier work this paper cites.
Model-free reinforcement learning as mixture learning
Vlassis, N. and Toussaint, M. (2009) · 2009
Earlier work this paper cites.
Double Q-learning
Hasselt, H. (2010) · 2010
Earlier work this paper cites.
Relative entropy policy search
Peters, J., Mulling, K., and Altun, Y. (2010) · 2010
Earlier work this paper cites.
Training and testing of recommender systems on data missing not at random
Steck, H. (2010) · 2010
Earlier work this paper cites.
Learning from logged implicit exploration data
Strehl, A., Langford, J., Li, L., and Kakade, S. M. (2010) · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Dudík, M., Langford, J., and Li, L. (2011) · 2011
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Bottou, L., Peters, J., Quiñonero-Candela, J., Charles, D. X., Chickering, D. M., Portugaly, E., Ray, D., Simard, P., and Snelson, E. (2013) · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Session-based recommendations with recurrent neural networks
Hidasi, B., Karatzoglou, A., Baltrunas, L., and Tikk, D. (2015) · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
Modeling user exposure in recommendation
Liang, D., Charlin, L., McInerney, J., and Blei, D. M. (2016) · 2016
Cited alongside, same era.
Recommendations as treatments: Debiasing learning and evaluation
Schnabel, T., Swaminathan, A., Singh, A., Chandak, N., and Joachims, T. (2016) · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E. (2016) · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017) · 2017
BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer
Sun, F., Liu, J., Wu, J., Pei, C., Lin, X., Ou, W., and Jiang, P. (2019) · 2019
Later among the works it cites.
On the design of estimators for bandit off-policy evaluation
Vlassis, N., Bibaut, A., Dimakopoulou, M., and Jebara, T. (2019) · 2019
Later among the works it cites.
A simple convolutional generative network for next item recommendation
Yuan, F., Karatzoglou, A., Arapakis, I., Jose, J. M., and He, X. (2019) · 2019
Later among the works it cites.
Reinforcement learning to optimize long-term user engagement in recommender systems
Zou, L., Xia, L., Ding, Z., Song, J., Liu, W., and Yin, D. (2019) · 2019
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S. (2020) · 2020
Later among the works it cites.
Off-policy bandits with deficient support
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Variational inference: A review for statisticians
Blei, D. M., Kucukelbir, A., and McAuliffe, J. D. (2017) · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Causal embeddings for recommendation
Bonner, S. and Vasile, F. (2018) · 2018
Cited alongside, same era.
Offline A/B testing for recommender systems
Gilotte, A., Calauzènes, C., Nedelec, T., Abraham, A., and Dollé, S. (2018) · 2018
Cited alongside, same era.
Self-attentive sequential recommendation
Kang, W.-C. and McAuley, J. (2018) · 2018
Cited alongside, same era.
Variational autoencoders for collaborative filtering
Liang, D., Krishnan, R. G., Hoffman, M. D., and Jebara, T. (2018) · 2018
Cited alongside, same era.
Sachdeva, N., Su, Y., and Joachims, T. (2020) · 2020
Later among the works it cites.
Causal inference for recommender systems
Wang, Y., Liang, D., Charlin, L., and Blei, D. M. (2020) · 2020
Later among the works it cites.
Self-supervised reinforcement learning for recommender systems
Xin, X., Karatzoglou, A., Arapakis, I., and Jose, J. M. (2020) · 2020
Later among the works it cites.
Offline RL without off-policy evaluation
Brandfonbrener, D., Whitney, W., Ranganath, R., and Bruna, J. (2021) · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. (2021) · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S. (2021) · 2021
Later among the works it cites.
Pessimistic reward models for off-policy learning in recommendation
Jeunen, O. and Goethals, B. (2021) · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. (2021) · 2021
Later among the works it cites.
Evaluating the robustness of off-policy evaluation
Saito, Y., Udagawa, T., Kiyohara, H., Mogi, K., Narita, Y., and Tateno, K. (2021) · 2021
Later among the works it cites.
Off-policy actor-critic for recommender systems
Chen, M., Xu, C., Gatto, V., Jain, D., Kumar, A., and Chi, E. (2022) · 2022
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D. (2019) · 2062
Closest in time.