Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) defines the task of learning from a static logged dataset without continually interacting with the environment.
Actor-Critic - Type Learning Algorithms for Markov Decision Processes
V. R. Konda and V. S. Borkar · 1999
Earlier work this paper cites.
Actor-Critic Algorithms
V. R. Konda and J. N. Tsitsiklis · 1999
Earlier work this paper cites.
Off-Policy Temporal Difference Learning with Function Approximation
D. Precup, R. S. Sutton, and S. Dasgupta · 2001
Earlier work this paper cites.
Approximately Optimal Approximate Reinforcement Learning
S. M. Kakade and J. Langford · 2002
Earlier work this paper cites.
Information Theory - Coding Theorems for Discrete Memoryless Systems, Second Edition
I. Csiszár and J. Körner · 2011
Earlier work this paper cites.
Batch Reinforcement Learning
S. Lange, T. Gabel, and M. A. Riedmiller · 2012
Earlier work this paper cites.
Auto-Encoding Variational Bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Generative Adversarial Nets
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Learning Structured Output Representation using Deep Conditional Generative Models
K. Sohn, H. Lee, and X. Yan · 2015
Earlier work this paper cites.
An Emphatic Approach to the Problem of Off-policy Temporal-Difference Learning
R. S. Sutton, A. R. Mahmood, and M. White · 2016
Earlier work this paper cites.
Constrained Policy Optimization
J. Achiam, D. Held, A. Tamar, and P. Abbeel · 2017
Earlier work this paper cites.
VEEGAN: Reducing Mode Collapse in GANs using Implicit Variational Learning
A. Srivastava, L. Valkov, C. Russell, M. U. Gutmann, and C. Sutton · 2017
Earlier work this paper cites.
Addressing Function Approximation Error in Actor-Critic Methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Earlier work this paper cites.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Deep Q-learning From Demonstrations
T. Hester, M. Vecerík, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, G. Dulac-Arnold, J. P. Agapiou, J. Z. Leibo, and A. Gruslys · 2018
Earlier work this paper cites.
Policy Optimization with Demonstrations
B. Kang, Z. Jie, and J. Feng · 2018
Earlier work this paper cites.
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, J. Schulman, E. Todorov, and S. Levine · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Exponentially Weighted Imitation Learning for Batched Historical Data
Q. Wang, J. Xiong, L. Han, P. Sun, H. Liu, and T. Zhang · 2018
Earlier work this paper cites.
Seeing What a GAN Cannot Generate
D. Bau, J.-Y. Zhu, J. Wulff, W. S. Peebles, H. Strobelt, B. Zhou, and A. Torralba · 2019
Earlier work this paper cites.
Off-Policy Deep Reinforcement Learning without Exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Earlier work this paper cites.
Off-Policy Deep Reinforcement Learning by Bootstrapping the Covariate Shift
C. Gelada and M. G. Bellemare · 2019
Earlier work this paper cites.
When to Trust Your Model: Model-Based Policy Optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Cited alongside, same era.
Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
A. Kumar, J. Fu, G. Tucker, and S. Levine · 2019
Cited alongside, same era.
Off-Policy Policy Gradient with State Distribution Correction
Y. Liu, A. Swaminathan, A. Agarwal, and E. Brunskill · 2019
Cited alongside, same era.
DualDICE: Behavior-Agnostic Estimation of Discounted Stationary Distribution Corrections
O. Nachum, Y. Chow, B. Dai, and L. Li · 2019
Cited alongside, same era.
Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
Y. Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. V. Dillon, B. Lakshminarayanan, and J. Snoek · 2019
Cited alongside, same era.
Offline Reinforcement Learning as One Big Sequence Modeling Problem
M. Janner, Q. Li, and S. Levine · 2021
Later among the works it cites.
Is Pessimism Provably Efficient for Offline RL?
Y. Jin, Z. Yang, and Z. Wang · 2021
Later among the works it cites.
Offline Reinforcement Learning with Fisher Divergence Critic Regularization
I. Kostrikov, J. Tompson, R. Fergus, and O. Nachum · 2021
Later among the works it cites.
Representation Balancing Offline Model-based Reinforcement Learning
B.-J. Lee, J. Lee, and K.-E. Kim · 2021
Later among the works it cites.
Offline-to-Online Reinforcement Learning via Balanced Replay and Pessimistic Q-Ensemble
S. Lee, Y. Seo, K. Lee, P. Abbeel, and J. Shin · 2021
Later among the works it cites.
Conservative Offline Distributional Reinforcement Learning
Y. J. Ma, D. Jayaraman, and O. Bastani · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Wu, G. Tucker, and O. Nachum · 2019
Cited alongside, same era.
BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement Learning
X. Chen, Z. Zhou, Z. Wang, C. Wang, Y. Wu, Q. Deng, and K. W. Ross · 2020
Cited alongside, same era.
D4RL: Datasets for Deep Data-Driven Reinforcement Learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
MOReL: Model-Based Offline Reinforcement Learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Cited alongside, same era.
Conservative Q-Learning for Offline Reinforcement Learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Accelerating Online Reinforcement Learning with Offline Datasets
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2020
Cited alongside, same era.
Critic Regularized Regression
Z. Wang, A. Novikov, K. Zolna, J. T. Springenberg, S. E. Reed, B. Shahriari, N. Siegel, J. Merel, C. Gulcehre, N. M. O. Heess, and N. de Freitas · 2020
Cited alongside, same era.
Later among the works it cites.
Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimization
T. Matsushima, H. Furuta, Y. Matsuo, O. Nachum, and S. Gu · 2021
Later among the works it cites.
Offline Pre-trained Multi-Agent Decision Transformer: One Big Sequence Model Tackles All SMAC Tasks
L. Meng, M. Wen, Y. Yang, C. Le, X. Li, W. Zhang, Y. Wen, H. Zhang, J. Wang, and B. Xu · 2021
Later among the works it cites.
Offline Reinforcement Learning from Images with Latent Space Models
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn · 2021
Later among the works it cites.
Bridging Offline reinforcement Learning and Imitation Learning: A Tale of Pessimism
P. Rashidinejad, B. Zhu, C. Ma, J. Jiao, and S. Russell · 2021
Later among the works it cites.
Offline Reinforcement Learning with Reverse Model-based Imagination
J. Wang, W. Li, H. Jiang, G. Zhu, S. Li, and C. Zhang · 2021
Later among the works it cites.
Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning
Y. Wu, S. Zhai, N. Srivastava, J. M. Susskind, J. Zhang, R. Salakhutdinov, and H. Goh · 2021
Later among the works it cites.
Bellman-Consistent Pessimism for Offline Reinforcement Learning
T. Xie, C.-A. Cheng, N. Jiang, P. Mineiro, and A. Agarwal · 2021
Later among the works it cites.
Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning
Y. Yang, X. Ma, L. Chenghao, Z. Zheng, Q. Zhang, G. Huang, J. Yang, and Q. Zhao · 2021
Later among the works it cites.
COMBO: Conservative Offline Model-Based Policy Optimization
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn · 2021
Later among the works it cites.
Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
A. Zanette, M. J. Wainwright, and E. Brunskill · 2021
Later among the works it cites.
Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning
C. Bai, L. Wang, Z. Yang, Z.-H. Deng, A. Garg, P. Liu, and Z. Wang · 2022
Closest in time.
Offline RL Policies Should be Trained to be Adaptive
D. Ghosh, A. Ajay, P. Agrawal, and S. Levine · 2022
Closest in time.
Offline Reinforcement Learning with Implicit Q-Learning
I. Kostrikov, A. Nair, and S. Levine · 2022
Closest in time.
Offline Reinforcement Learning with Value-based Episodic Memory
X. Ma, Y. Yang, H. Hu, J. Yang, C. Zhang, Q. Zhao, B. Liang, and Q. Liu · 2022
Closest in time.
A Regularized Implicit Policy for Offline Reinforcement Learning
S. Yang, Z. Wang, H. Zheng, Y. Feng, and M. Zhou · 2022
Closest in time.
Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement Learning, 2022
Y. Zhao, R. Boney, A. Ilin, J. Kannala, and J. Pajarinen · 2022
Closest in time.