Fetching the paper…
Reading the bibliography…
Offline reinforcement learning enables learning from a fixed dataset, without further interactions with the environment.
Divergence Measures Based on the Shannon Entropy
Lin, J · 1991
Earlier work this paper cites.
Self-Improving Reactive Agents Based on Reinforcement Learning, Planning and Teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Integral Probability Metrics and Their Generating Classes of Functions
Müller, A · 1997
Earlier work this paper cites.
The PageRank Citation Ranking: Bringing order to the Web
Page, L., Brin, S., Motwani, R., and Winograd, T · 1998
Earlier work this paper cites.
Matrix Analysis and Applied Linear Algebra
Meyer, C. D · 2000
Earlier work this paper cites.
Deeper inside PageRank
Langville, A. N. and Meyer, C. D · 2004
Earlier work this paper cites.
Tree-Based Batch Mode Reinforcement Learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
All of Nonparametric Statistics
Wasserman, L · 2006
Earlier work this paper cites.
Double Q-learning
Hasselt, H. V · 2010
Earlier work this paper cites.
T. E. Harris’s Contributions to Recurrent Markov Processes and Stochastic Flows
Baxendale, P · 2011
Earlier work this paper cites.
A Kernel Two-Sample Test
Gretton, A., Borgwardt, K., Rasch, M., Schölkopf, B., and Smola, A · 2012
Earlier work this paper cites.
Batch Reinforcement Learning , pp. 45–73
Lange, S., Gabel, T., and Riedmiller, M · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Equivalence of Distance-based and RKHS-based Statistics in Hypothesis Testing
Sejdinovic, D., Sriperumbudur, B., Gretton, A., and Fukumizu, K · 2013
Earlier work this paper cites.
Generative Adversarial Nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Conditional Generative Adversarial Nets
Mirza, M. and Osindero, S · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Trust Region Policy Optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P · 2015
Earlier work this paper cites.
Learning Structured Output Representation using Deep Conditional Generative Models
Sohn, K., Lee, H., and Yan, X · 2015
Earlier work this paper cites.
Nips 2016 tutorial: Generative adversarial networks
Goodfellow, I · 2016
Earlier work this paper cites.
Generative Adversarial Imitation Learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Continuous Control with Deep Reinforcement Learning
Lillicrap, T., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Earlier work this paper cites.
Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
Radford, A., Metz, L., and Chintala, S · 2016
Earlier work this paper cites.
Improved Techniques for Training GANs
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X · 2016
Cited alongside, same era.
White, T · 2016
Cited alongside, same era.
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Cited alongside, same era.
The Cramer Distance as a Solution to Biased Wasserstein Gradients
Bellemare, M. G., Danihelka, I., Dabney, W., Mohamed, S., Lakshminarayanan, B., Hoyer, S., and Munos, R · 2017
Cited alongside, same era.
Improved Training of Wasserstein GANs
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. C · 2017
Cited alongside, same era.
Learning When-to-Treat Policies
Nie, X., Brunskill, E., and Wager, S · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Confounding-robust policy evaluation in infinite-horizon reinforcement learning
Kallus, N. and Zhou, A · 2020
Later among the works it cites.
Conservative Q-Learning for Offline Reinforcement Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement Learning with Deep Energy-Based Policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
MMD GAN: Towards Deeper Understanding of Moment Matching Network
Li, C.-L., Chang, W.-C., Cheng, Y., Yang, Y., and Póczos, B · 2017
Cited alongside, same era.
Off-policy Evaluation for Slate Recommendation
Swaminathan, A., Krishnamurthy, A., Agarwal, A., Dudík, M., Langford, J., Jose, D., and Zitouni, I · 2017
Cited alongside, same era.
Deep Reinforcement Learning for Automated Radiation Adaptation in Lung Cancer
Tseng, H., Luo, Y., Cui, S., Chien, J.-T., Haken, R. T., and Naqa, I · 2017
Cited alongside, same era.
Binkowski, M., Sutherland, D. J., Arbel, M., and Gretton, A · 2018
Cited alongside, same era.
Many paths to equilibrium: GANs do not need to decrease a divergence at every step
Fedus, W., Rosca, M., Lakshminarayanan, B., Dai, A. M., Mohamed, S., and Goodfellow, I · 2018
Cited alongside, same era.
Addressing Function Approximation Error in Actor-Critic Methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile Critics
Kuznetsov, A., Shvechikov, P., Grishin, A., and Vetrov, D · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
Black-box Off-policy Estimation for Infinite-Horizon Reinforcement Learning
Mousavi, A., Li, L., Liu, Q., and Zhou, D · 2020
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Siegel, N. Y., Springenberg, J. T., Berkenkamp, F., Abdolmaleki, A., Neunert, M., Lampe, T., Hafner, R., Heess, N., and Riedmiller, M · 2020
Later among the works it cites.
Critic Regularized Regression
Wang, Z., Novikov, A., Zolna, K., Merel, J. S., Springenberg, J. T., Reed, S. E., Shahriari, B., Siegel, N., Gulcehre, C., Heess, N., and de Freitas, N · 2020
Later among the works it cites.
Implicit Distributional Reinforcement Learning
Yue, Y., Wang, Z., and Zhou, M · 2020
Later among the works it cites.
A Survey of Autonomous Driving: Common Practices and Emerging Technologies
Yurtsever, E., Lambert, J., Carballo, A., and Takeda, K · 2020
Later among the works it cites.
GenDICE: Generalized Offline Estimation of Stationary Values
Zhang, R., Dai, B., Li, L., and Schuurmans, D · 2020
Later among the works it cites.
Uncertainty-based offline reinforcement learning with diversified q-ensemble
An, G., Moon, S., Kim, J.-H., and Song, H. O · 2021
Later among the works it cites.
Behavioral Priors and Dynamics Models: Improving Performance and Domain Transfer in Offline RL
Cang, C., Rajeswaran, A., Abbeel, P., and Laskin, M · 2021
Later among the works it cites.
Decision Transformer: Reinforcement Learning via Sequence Modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Later among the works it cites.
A Minimalist Approach to Offline Reinforcement Learning
Fujimoto, S. and Gu, S. S · 2021
Later among the works it cites.
Addressing Extrapolation Error in Deep Offline Reinforcement Learning
Gulcehre, C., Colmenarejo, S. G., ziyu wang, Sygnowski, J., Paine, T., Zolna, K., Chen, Y., Hoffman, M., Pascanu, R., and de Freitas, N · 2021
Later among the works it cites.
Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimization
Matsushima, T., Furuta, H., Matsuo, Y., Nachum, O., and Gu, S. S · 2021
Later among the works it cites.
Risk-Averse Offline Reinforcement Learning
Urpí, N. A., Curi, S., and Krause, A · 2021
Later among the works it cites.
Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning
Wu, Y., Zhai, S., Srivastava, N., Susskind, J., Zhang, J., Salakhutdinov, R., and Goh, H · 2021
Later among the works it cites.
COMBO: Conservative Offline Model-Based Policy Optimization
Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C · 2021
Later among the works it cites.
S4rl: Surprisingly simple self-supervision for offline reinforcement learning in robotics
Sinha, S., Mandlekar, A., and Garg, A · 2022
Closest in time.