Fetching the paper…
Reading the bibliography…
We consider the offline reinforcement learning (RL) setting where the agent aims to optimize the policy solely from the data without further environment interactions.
AlgaeDICE: Policy gradient from arbitrary experience
Nachum, O., Dai, B., Kostrikov, I., Chow, Y., Li, L., and Schuurmans, D · 1912
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Convex optimization
Boyd, S., Boyd, S. P., and Vandenberghe, L · 2004
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Nilim, A. and El Ghaoui, L · 2005
Earlier work this paper cites.
The many faces of optimism: A unifying approach
Szita, I. and Lörincz, A · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Reinforcement learning: State-of-the-art
Lange, S., Gabel, T., and Riedmiller, M · 2012
Cited alongside, same era.
Generative adversarial networks
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Safe policy improvement by minimizing robust baseline regret
Petrik, M., Ghavamzadeh, M., and Chow, Y · 2016
Cited alongside, same era.
Equivalence between policy gradients and soft Q-learning, 2017
Schulman, J., Chen, X., and Abbeel, P · 2017
Cited alongside, same era.
Safe policy improvement with baseline bootstrapping
Laroche, R., Trichelair, P., and Des Combes, R. T · 2019
Later among the works it cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning, 2019
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning, 2019
Wu, Y., Tucker, G., and Nachum, O · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Later among the works it cites.
CoinDICE: Off-policy confidence interval estimation
Dai, B., Nachum, O., Chow, Y., Li, L., Szepesvari, C., and Schuurmans, D · 2020
Later among the works it cites.
MOReL : Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Cited alongside, same era.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog, 2019
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R · 2019
Cited alongside, same era.
Stabilizing off-policy Q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Cited alongside, same era.
Later among the works it cites.
Conservative Q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Batch reinforcement learning with hyperparameter gradients
Lee, B.-J., Lee, J., Vrancx, P., Kim, D., and Kim, K.-E · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
Off-policy evaluation via the regularized lagrangian
Yang, M., Nachum, O., Dai, B., Li, L., and Schuurmans, D · 2020
Later among the works it cites.
D4RL: Datasets for deep data-driven reinforcement learning, 2021
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2021
Closest in time.