Fetching the paper…
Reading the bibliography…
Offline Reinforcement Learning (RL) aims at learning an optimal control from a fixed dataset, without interactions with the system.
A possibility for implementing curiosity and boredom in model-building neural controllers
J. Schmidhuber · 1991
Earlier work this paper cites.
An approach to learning mobile robot navigation
S. Thrun · 1995
Earlier work this paper cites.
Biped dynamic walking using reinforcement learning
H. Benbrahim and J. A. Franklin · 1997
Earlier work this paper cites.
Autonomous helicopter control using reinforcement learning policy search methods
J. A. Bagnell and J. G. Schneider · 2001
Earlier work this paper cites.
Marginal mean models for dynamic regimes
S. A. Murphy, M. J. van der Laan, J. M. Robins, and C. P. P. R. Group · 2001
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2002
Earlier work this paper cites.
Least-squares policy iteration
M. G. Lagoudakis and R. Parr · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
D. Ernst, P. Geurts, and L. Wehenkel · 2005
Earlier work this paper cites.
New recommendation system using reinforcement learning
P. Rojanavasu, P. Srinil, and O. Pinngern · 2005
Earlier work this paper cites.
Learning cpg-based biped locomotion with a policy gradient method: Application to a humanoid robot
G. Endo, J. Morimoto, T. Matsubara, J. Nakanishi, and G. Cheng · 2008
Earlier work this paper cites.
Novelty or surprise?
A. Barto, M. Mirolli, and G. Baldassarre · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
J. Kober, J. A. Bagnell, and J. Peters · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Offline policy evaluation across representations with applications to educational games
T. Mandel, Y.-E. Liu, S. Levine, E. Brunskill, and Z. Popovic · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
K. Sohn, H. Lee, and X. Yan · 2015
Earlier work this paper cites.
Learning deep representations of appearance and motion for anomalous event detection
D. Xu, E. Ricci, Y. Yan, J. Song, and N. Sebe · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Surprise-based intrinsic motivation for deep reinforcement learning
J. Achiam and S. Sastry · 2017
Cited alongside, same era.
Count-based exploration with neural density models
G. Ostrovski, M. G. Bellemare, A. Oord, and R. Munos · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
A theory of regularized markov decision processes
M. Geist, B. Scherrer, and O. Pietquin · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, G. Tucker, and S. Levine · 2019
Later among the works it cites.
Safe policy improvement with baseline bootstrapping
R. Laroche, P. Trichelair, and R. T. Des Combes · 2019
Later among the works it cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep reinforcement learning framework for autonomous driving
A. E. Sallab, M. Abdou, E. Perot, and S. Yogamani · 2017
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
H. Tang, R. Houthooft, D. Foote, A. Stooke, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2017
Cited alongside, same era.
Anomaly detection with robust deep autoencoders
C. Zhou and R. C. Paffenroth · 2017
Cited alongside, same era.
Large-scale study of curiosity-driven learning
Y. Burda, H. Edwards, D. Pathak, A. Storkey, T. Darrell, and A. A. Efros · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl
S. K. S. Ghasemipour, D. Schuurmans, and S. S. Gu · 2020
Later among the works it cites.
A survey of deep learning techniques for autonomous driving
S. Grigorescu, B. Trasnea, T. Cocias, and G. Macesanu · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2020
Later among the works it cites.
Critic regularized regression
Z. Wang, A. Novikov, K. Żołna, J. T. Springenberg, S. Reed, B. Shahriari, N. Siegel, J. Merel, C. Gulcehre, N. Heess, et al · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
Reinforcement learning based recommender systems: A survey
M. M. Afsar, T. Crump, and B. Far · 2021
Closest in time.
The importance of pessimism in fixed-dataset policy optimization
J. Buckman, C. Gelada, and M. G. Bellemare · 2021
Closest in time.
Offline reinforcement learning with pseudometric learning
R. Dadashi, S. Rezaeifar, N. Vieillard, L. Hussenot, O. Pietquin, and M. Geist · 2021
Closest in time.
Regularized behavior value estimation
C. Gulcehre, S. G. Colmenarejo, Z. Wang, J. Sygnowski, T. Paine, K. Zolna, Y. Chen, M. Hoffman, R. Pascanu, and N. de Freitas · 2021
Closest in time.
Combo: Conservative offline model-based policy optimization
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn · 2021
Closest in time.