Fetching the paper…
Reading the bibliography…
The problem of offline reinforcement learning focuses on learning a good policy from a log of environment interactions.
Equivalence notions and model minimization in Markov decision processes
R. Givan, T. Dean, and M. Greig · 2003
Earlier work this paper cites.
Z. Wang, A. Novikov, K. Zolna, J. T. Springenberg, S. E. Reed, B. Shahriari, N. Y. Siegel, J. Merel, Ç. Gülçehre, N. Heess, and N. de Freitas · 2006
Earlier work this paper cites.
Learning invariant representations for reinforcement learning without reconstruction
A. Zhang, R. McAllister, R. Calandra, Y. Gal, and S. Levine · 2006
Earlier work this paper cites.
Bisimulation metrics for continuous markov decision processes
N. Ferns, P. Panangaden, and D. Precup · 2011
Earlier work this paper cites.
Metrics for finite markov decision processes
N. Ferns, P. Panangaden, and D. Precup · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
J. Kober, J. A. Bagnell, and J. Peters · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Xsede: accelerating scientific discovery
J. Towns, T. Cockerill, M. Dahan, I. Foster, K. Gaither, A. Grimshaw, V. Hazlewood, S. Lathrop, D. Lifka, G. D. Peterson, et al · 2014
Earlier work this paper cites.
Model-free episodic control, 2016
C. Blundell, B. Uria, A. Pritzel, Y. Li, A. Ruderman, J. Z. Leibo, J. Rae, D. Wierstra, and D. Hassabis · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
A. Pritzel, B. Uria, S. Srinivasan, A. P. Badia, O. Vinyals, D. Hassabis, D. Wierstra, and C. Blundell · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Scalable methods for computing state similarity in deterministic markov decision processes, 2019
P. S. Castro · 2019
Conservative q-learning for offline reinforcement learning, 2020
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
Exact (then approximate) dynamic programming for deep reinforcement learning
H. Marklund, S. Nair, and C. Finn · 2020
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization, 2020
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
Episodic reinforcement learning with associative memory
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
When to trust your model: Model-based policy optimization, 2019
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction, 2019
A. Kumar, J. Fu, G. Tucker, and S. Levine · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning, 2019
Y. Wu, G. Tucker, and O. Nachum · 2019
Cited alongside, same era.
Bail: Best-action imitation learning for batch deep reinforcement learning, 2020
X. Chen, Z. Zhou, Z. Wang, C. Wang, Y. Wu, and K. Ross · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Cited alongside, same era.
G. Zhu, Z. Lin, G. Yang, and C. Zhang · 2020
Later among the works it cites.
Generalizable episodic memory for deep reinforcement learning
H. Hu, J. Ye, Z. Ren, G. Zhu, and C. Zhang · 2021
Later among the works it cites.
Deepaveragers: Offline reinforcement learning by solving derived non-parametric {mdp}s
A. K. Shrestha, S. Lee, P. Tadepalli, and A. Fern · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization, 2021
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn · 2021
Later among the works it cites.