Fetching the paper…
Reading the bibliography…
Reinforcement learning based recommender systems (RL-based RS) aim at learning a good policy from a batch of collected data, by casting recommendations to multi-step decision-making tasks.
Recsim: A configurable simulation platform for recommender systems
E. Ie, C.-w. Hsu, M. Mladenov, V. Jain, S. Narvekar, J. Wang, R. Wu, and C. Boutilier · 1909
Earlier work this paper cites.
A generalization of sampling without replacement from a finite universe
D. G. Horvitz and D. J. Thompson · 1952
Earlier work this paper cites.
The md5 message-digest algorithm, 1992
R. Rivest and S. Dusse · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
An mdp-based recommender system
G. Shani, D. Heckerman, R. I. Brafman, and C. Boutilier · 2005
Earlier work this paper cites.
Doubly robust policy evaluation and learning
M. Dudík, J. Langford, and L. Li · 2011
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
G. Dulac-Arnold, R. Evans, H. van Hasselt, P. Sunehag, T. Lillicrap, J. Hunt, T. Mann, T. Weber, T. Degris, and B. Coppin · 2015
Earlier work this paper cites.
Session-based recommendations with recurrent neural networks
B. Hidasi, A. Karatzoglou, L. Baltrunas, and D. Tikk · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Wide & deep learning for recommender systems
H.-T. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir, et al · 2016
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Data-efficient off-policy policy evaluation for reinforcement learning
P. Thomas and E. Brunskill · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Deep reinforcement learning for list-wise recommendations
X. Zhao, L. Zhang, L. Xia, Z. Ding, D. Yin, and J. Tang · 2017
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Cited alongside, same era.
Horizon: Facebook’s open source applied reinforcement learning platform
J. Gauci, E. Conti, Y. Liang, K. Virochsiri, Y. R. He, Z. Kaden, V. Narayanan, X. Ye, and S. Fujimoto · 2018
Cited alongside, same era.
Model-based reinforcement learning with adversarial training for online recommendation
X. Bai, J. Guan, and H. Wang · 2019
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Later among the works it cites.
Exact-k recommendation via maximal clique optimization
Y. Gong, Y. Zhu, L. Duan, Q. Liu, Z. Guan, F. Sun, W. Ou, and K. Q. Zhu · 2019
Later among the works it cites.
Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning
J. Shi, Y. Yu, Q. Da, S.-Y. Chen, and A. Zeng · 2019
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Rl unplugged: Benchmarks for offline reinforcement learning
C. Gulcehre, Z. Wang, A. Novikov, T. L. Paine, and N. D. Freitas · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. Daumé III, and K. Crawford · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Cited alongside, same era.
Reinforcement learning to rank in e-commerce search engine: Formalization, analysis, and application
Y. Hu, Q. Da, A. Zeng, Y. Yu, and Y. Xu · 2018
Cited alongside, same era.
RLlib: Abstractions for distributed reinforcement learning
E. Liang, R. Liaw, R. Nishihara, P. Moritz, R. Fox, K. Goldberg, J. E. Gonzalez, M. I. Jordan, and I. Stoica · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
A. Radford and K. Narasimhan · 2018
Cited alongside, same era.
D. Rohde, S. Bonner, T. Dunlop, F. Vasile, and A. Karatzoglou · 2018
Cited alongside, same era.
Deepmind control suite
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. Casas, D. Budden, A. Abdolmaleki, J. Merel, and A. Lefrancq · 2018
Cited alongside, same era.
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
d3rlpy: An offline deep reinforcement library
T. Seno · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, L. Kevin, R. Aravind, L. Kimin, G. Aditya, L. Michael, A. Pieter, S. Aravind, and M. Igor · 2021
Closest in time.
Aliexpress learning-to-rank: Maximizing online model performance without going online
G. Huzhang, Z. Pang, Y. Gao, Y. Liu, W. Shen, W.-J. Zhou, Q. Da, A. Zeng, H. Yu, Y. Yu, et al · 2021
Closest in time.
Reinforcement learning as one big sequence modeling problem
M. Janner, Q. Li, and S. Levine · 2021
Closest in time.
Accordion: a trainable simulator for long-term interactive systems
J. McInerney, E. Elahi, J. Basilico, Y. Raimond, and T. Jebara · 2021
Closest in time.
Combo: Conservative offline model-based policy optimization
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn · 2021
Closest in time.