Fetching the paper…
Reading the bibliography…
Offline RL algorithms must account for the fact that the dataset they are provided may leave many facets of the environment unknown.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Learning without state-estimation in partially observable markovian decision processes
Singh, S. P., Jaakkola, T. S., and Jordan, M. I · 1994
Earlier work this paper cites.
Optimal Learning: Computational procedures for Bayes-adaptive Markov decision processes
Duff, M. O · 2002
Earlier work this paper cites.
A survey of partially observable markov decision processes: Theory, models, and algorithms
Monahan, G. E · 2007
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Ramachandran, D. and Amir, E · 2007
Earlier work this paper cites.
Bayesian multi-task reinforcement learning
Lazaric, A. and Ghavamzadeh, M · 2010
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M. A · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Roy, B. V · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Bayesian reinforcement learning: A survey
Ghavamzadeh, M., Mannor, S., Pineau, J., and Tamar, A · 2015
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hasselt, H. V., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H. V., and Meger, D · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
A bayesian approach to generative adversarial imitation learning
Jeon, W., Seo, S., and Kim, K.-E · 2018
Cited alongside, same era.
If maxent rl is the answer, what is the question?
Eysenbach, B. and Levine, S · 2019
Cited alongside, same era.
Varibad: A very good method for bayes-adaptive deep rl via meta-learning
Zintgraf, L. M., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S · 2020
Later among the works it cites.
Uncertainty-based offline reinforcement learning with diversified q-ensemble
An, G., Moon, S., Kim, J.-H., and Song, H. O · 2021
Later among the works it cites.
Randomized ensembled double q-learning: Learning fast without a model
Chen, X., Wang, C., Zhou, Z., and Ross, K. W · 2021
Later among the works it cites.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl
Ghasemipour, S. K. S., Schuurmans, D., and Gu, S. S · 2021
Later among the works it cites.
Why generalization in rl is difficult: Epistemic pomdps and implicit partial observability
Ghosh, D., Rahme, J., Kumar, A., Zhang, A., Adams, R. P., and Levine, S · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Tucker, G., and Levine, S · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Cited alongside, same era.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Cited alongside, same era.
Offline meta learning of exploration
Dorfman, R. and Tamar, A · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Jin, Y., Yang, Z., and Wang, Z · 2021
Later among the works it cites.
A workflow for offline model-free robotic reinforcement learning
Kumar, A., Singh, A., Tian, S., Finn, C., and Levine, S · 2021
Later among the works it cites.
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
Lee, K., Laskin, M., Srinivas, A., and Abbeel, P · 2021
Later among the works it cites.
Mesa: Offline meta-rl for safe adaptation and fault tolerance
Luo, M., Balakrishna, A., Thananjeyan, B., Nair, S., Ibarz, J., Tan, J., Finn, C., Stoica, I., and Goldberg, K · 2021
Later among the works it cites.
What matters in learning from offline human demonstrations for robot manipulation
Mandlekar, A., Xu, D., Wong, J., Nasiriany, S., Wang, C., Kulkarni, R., Fei-Fei, L., Savarese, S., Zhu, Y., and Martín-Martín, R · 2021
Later among the works it cites.
Ghasemipour, S. K. S., Gu, S. S., and Nachum, O · 2022
Closest in time.
Discrete-Time Birth-Death Chains, aug 10 2020
Siegrist, K · 2022
Closest in time.