Fetching the paper…
Reading the bibliography…
Batch Reinforcement Learning (Batch RL) consists in training a policy using trajectories collected with another policy, called the behavioural policy.
Linear programming and extensions
Dantzig, G. (1963) · 1963
Earlier work this paper cites.
How good is the simplex algorithm?
Klee, V. and Minty, G. J. (1972) · 1972
Earlier work this paper cites.
The Simplex Method: A Probabilistic Analysis
Borgwardt, K. H. (1987) · 1987
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Reinforcement learning for spoken dialogue systems
Singh, S. P., Kearns, M. J., Litman, D. J., and Walker, M. A. (1999) · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
Linear Programming 2: Theory and Extensions
Dantzig, G. B. and Thapa, M. N. (2003) · 2003
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Weissman, T., Ordentlich, E., Seroussi, G., Verdu, S., and Weinberger, M. J. (2003) · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L. (2005) · 2005
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N. (2005) · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Nilim, A. and El Ghaoui, L. (2005) · 2005
Cited alongside, same era.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M. (2005) · 2005
Cited alongside, same era.
Adaptive treatment of epilepsy via batch-mode reinforcement learning
Guez, A., Vincent, R. D., Avoli, M., and Pineau, J. (2008) · 2008
Cited alongside, same era.
Interior point methods 25 years later
Gondzio, J. (2012) · 2012
Cited alongside, same era.
Batch Reinforcement Learning
Lange, S., Gabel, T., and Riedmiller, M. (2012) · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G. (2012) · 2012
Cited alongside, same era.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D. (2015) · 2015
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. (2016) · 2016
Later among the works it cites.
Safe policy improvement by minimizing robust baseline regret
Petrik, M., Ghavamzadeh, M., and Chow, Y. (2016) · 2016
Later among the works it cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017) · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
Safe reinforcement learning
Thomas, P. S. (2015) · 2015
Cited alongside, same era.
Dora the explorer: Directed outreaching reinforcement action-selection
Fox, L., Choshen, L., and Loewenstein, Y. (2018) · 2018
Later among the works it cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O. (2019) · 2019
Closest in time.
A theory of regularized markov decision processes
Geist, M., Scherrer, B., and Pietquin, O. (2019) · 2019
Closest in time.
Safe policy improvement with baseline bootstrapping
Laroche, R., Trichelair, P., and Tachet des Combes, R. (2019) · 2019
Closest in time.
Safe policy improvement with baseline bootstrapping in factored environments
Simão, T. D. and Spaan, M. T. J. (2019) · 2019
Closest in time.