Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) aims at learning a good policy from a batch of collected data, without extra interactions with the environment during training.
Reinforcement Learning - An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder P. Singh · 2000
Earlier work this paper cites.
Statistical comparisons of classifiers over multi- ple data sets
J. Demšar · 2006
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin A. Riedmiller · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, and et al · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, and et al · 2015
Earlier work this paper cites.
CARLA: An open urban driving simulator
Alexey Dosovitskiy, Germán Ros, Felipe Codevilla, and et al · 2017
Earlier work this paper cites.
A benchmark environment motivated by industrial control problems
Daniel Hein, Stefan Depeweg, Michel Tokic, Steffen Udluft, Alexander Hentschel, Thomas A. Runkler, and Volkmar Sterzing · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, and et al · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, and et al · 2018
Cited alongside, same era.
Benchmarking batch deep reinforcement learning algorithms
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and et al · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Abhishek Gupta, Vikash Kumar, Corey Lynch, and et al · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Later among the works it cites.
An empirical investigation of the challenges of real-world reinforcement learning
Gabriel Dulac-Arnold, Nir Levine, Daniel J. Mankowitz, Jerry Li, Cosmin Paduraru, Sven Gowal, and Todd Hester · 2020
Later among the works it cites.
D4RL: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, and et al · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stabilizing off-policy Q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, and et al · 2019
Cited alongside, same era.
Batch policy learning under constraints
Hoang Minh Le, Cameron Voloshin, and Yisong Yue · 2019
Cited alongside, same era.
Environment reconstruction with hidden confounders for reinforcement learning based recommendation
Wenjie Shang, Yang Yu, Qingyang Li, Zhiwei Qin, Yiping Meng, and Jieping Ye · 2019
Cited alongside, same era.
Virtual-Taobao: Virtualizing real-world online retail environment for reinforcement learning
Jing-Cheng Shi, Yang Yu, Qing Da, Shi-Yong Chen, and An-Xiang Zeng · 2019
Cited alongside, same era.
Empirical study of off-policy policy evaluation for reinforcement learning
Cameron Voloshin, Hoang Minh Le, Nan Jiang, and Yisong Yue · 2019
Cited alongside, same era.
CityLearn v1.0: An OpenAI Gym environment for demand response with deep reinforcement learning
José R. Vázquez-Canteli, Jérôme Kämpf, Gregor Henze, and Zoltan Nagy · 2019
Cited alongside, same era.
RL unplugged: A collection of benchmarks for offline reinforcement learning
Çaglar Gülçehre, Ziyu Wang, Alexander Novikov, and et al
Cited in the paper.
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
FinRL: A deep reinforcement learning library for automated stock trading in quantitative finance
Xiao-Yang Liu, Hongyang Yang, Qian Chen, Runjia Zhang, Liuqing Yang, Bowen Xiao, and Christina Dan Wang · 2020
Later among the works it cites.
Hyperparameter selection for offline reinforcement learning
Tom Le Paine, Cosmin Paduraru, Andrea Michi, and et al · 2020
Later among the works it cites.
Error bounds of imitating policies and environments
Tian Xu, Ziniu Li, and Yang Yu · 2020
Later among the works it cites.
MOPO: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, and et al · 2020
Later among the works it cites.
PLAS: Latent action space for offline reinforcement learning
Wenxuan Zhou, Sujay Bajracharya, and David Held · 2020
Later among the works it cites.