Fetching the paper…
Reading the bibliography…
This paper considers Safe Policy Improvement (SPI) in Batch Reinforcement Learning (Batch RL): from a fixed dataset and without direct access to the true environment, train a policy that is guaranteed to perform at least as well as the baseline policy used to collect the data.
Scaling up budgeted reinforcement learning
Carrara, N., Leurent, E., Laroche, R., Urvoy, T., Maillard, O., and Pietquin, O · 1903
Earlier work this paper cites.
A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D · 1957
Earlier work this paper cites.
Dynamic programming
Howard, R. A · 1966
Earlier work this paper cites.
On the convergence of policy iteration in stationary dynamic programming
Puterman, M. L. and Brumelle, S. L · 1979
Earlier work this paper cites.
Better bootstrap confidence intervals
Efron, B · 1987
Earlier work this paper cites.
Hierarchical control and learning for Markov decision processes
Parr, R. E. and Russell, S · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Constrained Markov Decision Processes
Altman, E · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
A reinforcement learning algorithm based on policy iteration for average reward: Empirical results with yield management and convergence analysis
Gosavi, A · 2004
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Nilim, A. and El Ghaoui, L · 2005
Cited alongside, same era.
Bandit based monte-carlo planning
Kocsis, L. and Szepesvári, C · 2006
Cited alongside, same era.
Knows what it knows: a framework for self-aware learning
Li, L., Littman, M. L., and Walsh, T. J · 2008
Cited alongside, same era.
The many faces of optimism: a unifying approach
Szita, I. and Lőrincz, A · 2008
Cited alongside, same era.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P · 2009
Cited alongside, same era.
Percentile optimization for markov decision processes with parameter uncertainty
Delage, E. and Mannor, S · 2010
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Later among the works it cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimising a handcrafted dialogue system design
Laroche, R., Putois, G., and Bretier, P · 2010
Cited alongside, same era.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
Off-policy Evaluation in Markov Decision Processes
Paduraru, C · 2013
Cited alongside, same era.
Offline policy evaluation across representations with applications to educational games
Mandel, T., Liu, Y.-E., Levine, S., Brunskill, E., and Popovic, Z · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Cited alongside, same era.
Safely interruptible agents
Orseau, L. and Armstrong, S · 2016
Later among the works it cites.
Safe policy improvement by minimizing robust baseline regret
Petrik, M., Ghavamzadeh, M., and Chow, Y · 2016
Later among the works it cites.
Dynamic safe interruptibility for decentralized multi-agent reinforcement learning
Guerraoui, R., Hendrikx, H., Maurer, A., et al · 2017
Closest in time.
Using options and covariance testing for long horizon off-policy policy evaluation
Guo, Z., Thomas, P. S., and Brunskill, E · 2017
Closest in time.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Closest in time.
On ensuring that intelligent machines are well-behaved
Thomas, P. S., da Silva, B. C., Barto, A. G., and Brunskill, E · 2017
Closest in time.
Dora the explorer: Directed outreaching reinforcement action-selection
Fox, L., Choshen, L., and Loewenstein, Y · 2018
Closest in time.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2019
Closest in time.
Dead-ends and secure exploration in reinforcement learning
Fatemi, M., Sharma, Shikhar and van Seijen, H., and Ebrahimi Kahou, S · 2019
Closest in time.