Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL), also known as batch RL, aims to optimize policy from a large pre-recorded dataset without interaction with the environment.
L.-J. Lin, “Self-improving reactive agents based on reinforcement learning, planning and teaching,” Machine learning , vol. 8, no. 3-4, pp. 293–321, 1992
1992
Earlier work this paper cites.
G. J. Gordon, “Stable function approximation in dynamic programming,” in Machine Learning Proceedings 1995 . Elsevier, 1995, pp. 261–268
1995
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Advances in neural information processing systems , 2000, pp. 1057–1063
2000
Earlier work this paper cites.
D. Ormoneit and Ś. Sen, “Kernel-based reinforcement learning,” Machine learning , vol. 49, no. 2-3, pp. 161–178, 2002
2002
Earlier work this paper cites.
D. Ernst, P. Geurts, and L. Wehenkel, “Tree-based batch mode reinforcement learning,” Journal of Machine Learning Research , vol. 6, no. Apr, pp. 503–556, 2005
2005
Earlier work this paper cites.
M. Riedmiller, “Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method,” in European Conference on Machine Learning . Springer, 2005, pp. 317–328
2005
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing data using t-sne,” Journal of machine learning research , vol. 9, no. Nov, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola, “A kernel two-sample test,” The Journal of Machine Learning Research , vol. 13, no. 1, pp. 723–773, 2012
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
B. Kim, A.-m. Farahmand, J. Pineau, and D. Precup, “Learning from limited demonstrations,” in Advances in Neural Information Processing Systems , 2013, pp. 2859–2867
2013
Earlier work this paper cites.
B. Piot, M. Geist, and O. Pietquin, “Boosted bellman residual minimization handling expert demonstrations,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 2014, pp. 549–564
2014
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” 2014
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
J. Chemali and A. Lazaric, “Direct policy iteration with demonstrations,” in IJCAI-24th International Joint Conference on Artificial Intelligence , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
2016
Cited alongside, same era.
J. Ho, J. Gupta, and S. Ermon, “Model-free imitation learning with policy optimization,” in International Conference on Machine Learning , 2016, pp. 2760–2769
2016
Cited alongside, same era.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in Thirtieth AAAI conference on artificial intelligence , 2016
2016
Cited alongside, same era.
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al. , “Grandmaster level in starcraft ii using multi-agent reinforcement learning,” Nature , vol. 575, no. 7782, pp. 350–354, 2019
2019
Later among the works it cites.
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine, “Stabilizing off-policy q-learning via bootstrapping error reduction,” in Advances in Neural Information Processing Systems , 2019, pp. 11 784–11 794
2019
Later among the works it cites.
R. Laroche, P. Trichelair, and R. T. Des Combes, “Safe policy improvement with baseline bootstrapping,” in International Conference on Machine Learning , 2019, pp. 3652–3661
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al. , “A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,” Science , vol. 362, no. 6419, pp. 1140–1144, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Overcoming exploration in reinforcement learning with demonstrations,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 6292–6299
2018
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
O. Nachum, Y. Chow, B. Dai, and L. Li, “Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections,” in Advances in Neural Information Processing Systems , 2019, pp. 2318–2328
2019
Later among the works it cites.
2019
Later among the works it cites.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
S. Fujimoto, D. Meger, and D. Precup, “Off-policy deep reinforcement learning without exploration,” in International Conference on Machine Learning , 2019, pp. 2052–2062
2062
Closest in time.