Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) Algorithms are often designed with environments such as MuJoCo in mind, in which the planning horizon is extremely long and no noise exists.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Nachum, O., Chow, Y., Dai, B., and Li, L. (2019a) · 1906
Earlier work this paper cites.
BAIL: Best-action imitation learning for batch deep reinforcement learning
Chen, X., Zhou, Z., Wang, Z., Wang, C., Wu, Y., and Ross, K. (2019) · 1910
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O. (2019) · 1911
Earlier work this paper cites.
Algaedice: Policy gradient from arbitrary experience
Nachum, O., Dai, B., Kostrikov, I., Chow, Y., Li, L., and Schuurmans, D. (2019b) · 1912
Earlier work this paper cites.
A natural policy gradient
Kakade, S.M. (2001) · 2001
Earlier work this paper cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Siegel, N.Y., Springenberg, J.T., Berkenkamp, F., Abdolmaleki, A., Neunert, M., Lampe, T., Hafner, R., and Riedmiller, M. (2020) · 2002
Earlier work this paper cites.
Gendice: Generalized offline estimation of stationary values
Zhang, R., Dai, B., Li, L., and Schuurmans, D. (2020a) · 2002
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M.G. and Parr, R. (2003) · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L. (2005) · 2005
Earlier work this paper cites.
MOReL: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T. (2020) · 2005
Earlier work this paper cites.
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M. (2005) · 2005
Earlier work this paper cites.
Conservative Q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S. (2020) · 2006
Earlier work this paper cites.
Wang, Z., Novikov, A., Zolna, K., Springenberg, J.T., Reed, S., Shahriari, B., Siegel, N., Merel, J., Gulcehre, C., Heess, N., et al. (2020) · 2006
Earlier work this paper cites.
Provably good batch reinforcement learning without great exploration
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E. (2020) · 2007
Earlier work this paper cites.
Hyperparameter selection for offline reinforcement learning
Paine, T.L., Paduraru, C., Michi, A., Gulcehre, C., Zolna, K., Novikov, A., Wang, Z., and de Freitas, N. (2020) · 2007
Earlier work this paper cites.
A recurrent control neural network for data efficient reinforcement learning
Schaefer, A.M., Udluft, S., and Zimmermann, H.G. (2007) · 2007
Cited alongside, same era.
Improving optimality of neural rewards regression for data-efficient batch near-optimal policy identification
Schneegaß, D., Udluft, S., and Martinetz, T. (2007) · 2007
Cited alongside, same era.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
Ajay, A., Kumar, A., Agrawal, P., Levine, S., and Nachum, O. (2020) · 2010
Cited alongside, same era.
Deepaveragers: Offline reinforcement learning by solving derived non-parametric MDPs
Shrestha, A., Lee, S., Tadepalli, P., and Fern, A. (2020) · 2010
Cited alongside, same era.
Agent self-assessment: Determining policy quality without execution
Hans, A., Duell, S., and Udluft, S. (2011) · 2011
Cited alongside, same era.
Interpretable policies for reinforcement learning by genetic programming
Hein, D., Udluft, S., and Runkler, T.A. (2018) · 2018
Later among the works it cites.
Stabilizing off-policy Q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S. (2019) · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M. (2020) · 2020
Later among the works it cites.
Bayesian decomposition of multi-modal dynamical systems for reinforcement learning
Kaiser, M., Otte, C., Runkler, T.A., and Ek, C.H. (2020) · 2020
Later among the works it cites.
Reducing sampling error in batch temporal difference learning
Pavse, B., Durugkar, I., Hanna, J., and Stone, P. (2020) · 2020
Later among the works it cites.
MOPO: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J.Y., Levine, S., Finn, C., and Ma, T. (2020) · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
MuJoCo: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
Learning and policy search in stochastic dynamical systems with Bayesian neural networks
Depeweg, S., Hernández-Lobato, J.M., Doshi-Velez, F., and Udluft, S. (2016) · 2016
Cited alongside, same era.
Reinforcement learning with particle swarm optimization policy (PSO-P) in continuous state and action spaces
Hein, D., Hentschel, A., Runkler, T.A., and Udluft, S. (2016) · 2016
Cited alongside, same era.
Deep reinforcement learning with double Q-learning
Van Hasselt, H., Guez, A., and Silver, D. (2016) · 2016
Cited alongside, same era.
Decomposition of uncertainty in Bayesian deep learning for efficient and risk-sensitive learning
Depeweg, S., Hernández-Lobato, J.M., Doshi-Velez, F., and Udluft, S. (2017) · 2017
Cited alongside, same era.
A benchmark environment motivated by industrial control problems
Hein, D., Depeweg, S., Tokic, M., Udluft, S., Hentschel, A., Runkler, T.A., and Sterzing, V. (2017) · 2017
Cited alongside, same era.
Later among the works it cites.
Benchmarks for deep off-policy evaluation
Fu, J., Norouzi, M., Nachum, O., Tucker, G., Wang, Z., Novikov, A., Yang, M., Zhang, M.R., Chen, Y., Kumar, A., et al. (2021) · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S.S. (2021) · 2021
Later among the works it cites.
Emaq: Expected-max Q-learning operator for simple yet effective offline and online RL
Ghasemipour, S.K.S., Schuurmans, D., and Gu, S.S. (2021) · 2021
Later among the works it cites.
Active offline policy selection
Konyushkova, K., Chen, Y., Paine, T., Gulcehre, C., Paduraru, C., Mankowitz, D.J., Denil, M., and de Freitas, N. (2021) · 2021
Later among the works it cites.
Curriculum offline imitating learning
Liu, M., Zhao, H., Yang, Z., Shen, J., Zhang, W., Zhao, L., and Liu, T.Y. (2021) · 2021
Later among the works it cites.
PEBL: Pessimistic ensembles for offline deep reinforcement learning
Smit, J., Ponnambalam, C.T., Spaan, M.T., and Oliehoek, F.A. (2021) · 2021
Later among the works it cites.
Risk-averse offline reinforcement learning
Urpí, N.A., Curi, S., and Krause, A. (2021) · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization
Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C. (2021) · 2021
Later among the works it cites.
Autoregressive dynamics models for offline policy evaluation and optimization
Zhang, M.R., Paine, T.L., Nachum, O., Paduraru, C., Tucker, G., Wang, Z., and Norouzi, M. (2021) · 2021
Later among the works it cites.