Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) has increasingly become the focus of the artificial intelligent research due to its wide real-world applications where the collection of data may be difficult, time-consuming, or costly.
Benchmarking batch deep reinforcement learning algorithms
Fujimoto, S., Conti, E., Ghavamzadeh, M., and Pineau, J · 1910
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Weissman, T., Ordentlich, E., Seroussi, G., Verdu, S., and Weinberger, M. J · 2003
Earlier work this paper cites.
Neural network ensembles in reinforcement learning
Faußer, S. and Schwenker, F · 2015
Earlier work this paper cites.
Safe policy improvement by minimizing robust baseline regret
Ghavamzadeh, M., Petrik, M., and Chow, Y · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Earlier work this paper cites.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Anschel, O., Baram, N., and Shimkin, N · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Distributional reinforcement learning with quantile regression
Dabney, W., Rowland, M., Bellemare, M. G., and Munos, R · 2017
Cited alongside, same era.
Robust imitation of diverse behaviors
Wang, Z., Merel, J., Reed, S., Wayne, G., de Freitas, N., and Heess, N · 2017
Cited alongside, same era.
Reinforcement mechanism design for e-commerce
Cai, Q., Filos-Ratsikas, A., Tang, P., and Zhang, Y · 2018
Cited alongside, same era.
Reinforcement learning to rank in e-commerce search engine: Formalization, analysis, and application
Hu, Y., Da, Q., Zeng, A., Yu, Y., and Xu, Y · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Later among the works it cites.
Safe policy improvement with baseline bootstrapping
Laroche, R., Trichelair, P., and Des Combes, R. T · 2019
Later among the works it cites.
Model-based constrained mdp for budget allocation in sequential incentive marketing
Xiao, S., Guo, L., Jiang, Z., Lv, L., Chen, Y., Zhu, J., and Yang, S · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Later among the works it cites.
Bail: Best-action imitation learning for batch deep reinforcement learning
Chen, X., Zhou, Z., Wang, Z., Wang, C., Wu, Y., and Ross, K · 2020
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2062
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhao, X., Xia, L., Zhang, L., Ding, Z., Yin, D., and Tang, J · 2018
Cited alongside, same era.
Closest in time.