Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) shows promise of applying RL to real-world problems by effectively utilizing previously collected data.
Asymmetric least squares estimation and testing
Whitney K Newey and James L Powell · 1987
Earlier work this paper cites.
GenDICE: Generalized offline estimation of stationary values
Ruiyi Zhang, Bo Dai, Lihong Li, and Dale Schuurmans · 2002
Earlier work this paper cites.
Charles Blundell, Benigno Uria, Alexander Pritzel, Yazhe Li, Avraham Ruderman, Joel Z Leibo, Jack Rae, Daan Wierstra, and Demis Hassabis · 2016
Earlier work this paper cites.
Neural episodic control
Alexander Pritzel, Benigno Uria, Sriram Srinivasan, Adria Puigdomenech Badia, Oriol Vinyals, Demis Hassabis, Daan Wierstra, and Charles Blundell · 2017
Earlier work this paper cites.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc G Bellemare, and Rémi Munos · 2018
Earlier work this paper cites.
Episodic memory deep Q-networks
Zichuan Lin, Tianqi Zhao, Guangwen Yang, and Lintao Zhang · 2018
Earlier work this paper cites.
Exponentially weighted imitation learning for batched historical data
Qing Wang, Jiechao Xiong, Lei Han, Peng Sun, Han Liu, and Tong Zhang · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Earlier work this paper cites.
Stabilizing off-policy Q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Earlier work this paper cites.
Statistics and samples in distributional reinforcement learning
Mark Rowland, Robert Dadashi, Saurabh Kumar, Rémi Munos, Marc G Bellemare, and Will Dabney · 2019
Cited alongside, same era.
Arthur Argenson and Gabriel Dulac-Arnold · 2020
Cited alongside, same era.
BAIL: Best-action imitation learning for batch deep reinforcement learning
Xinyue Chen, Zijian Zhou, Zheng Wang, Che Wang, Yanqiu Wu, and Keith Ross · 2020
Cited alongside, same era.
D4RL: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
MOPO: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Offline reinforcement learning with pseudometric learning
Robert Dadashi, Shideh Rezaeifar, Nino Vieillard, Léonard Hussenot, Olivier Pietquin, and Matthieu Geist · 2021
Closest in time.
EMaQ: Expected-max Q-learning operator for simple yet effective offline and online RL
Seyed Kamyar Seyed Ghasemipour, Dale Schuurmans, and Shixiang Shane Gu · 2021
Closest in time.
Generalizable episodic memory for deep reinforcement learning
Hao Hu, Jianing Ye, Zhizhou Ren, Guangxiang Zhu, and Chongjie Zhang · 2021
Closest in time.
Is pessimism provably efficient for offline RL?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Conservative Q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2020
Cited alongside, same era.
Adaptive trade-offs in off-policy learning
Mark Rowland, Will Dabney, and Rémi Munos · 2020
Cited alongside, same era.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Noah Y Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller · 2020
Cited alongside, same era.
Self-imitation learning via generalized lower bound Q-learning
Yunhao Tang · 2020
Cited alongside, same era.
GradientDICE: Rethinking generalized offline estimation of stationary values
Shangtong Zhang, Bo Liu, and Shimon Whiteson
Cited in the paper.
Offline reinforcement learning with fisher divergence critic regularization
Ilya Kostrikov, Rob Fergus, Jonathan Tompson, and Ofir Nachum · 2021
Closest in time.
OptiDICE: Offline policy optimization via stationary distribution correction estimation
Jongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau, and Kee-Eung Kim · 2021
Closest in time.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Closest in time.
Instabilities of offline rl with pre-trained neural representation
Ruosong Wang, Yifan Wu, Ruslan Salakhutdinov, and Sham M Kakade · 2021
Closest in time.
Uncertainty weighted Actor-Critic for offline reinforcement learning
Yue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua Susskind, Jian Zhang, Ruslan Salakhutdinov, and Hanlin Goh · 2021
Closest in time.
Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning
Yiqin Yang, Xiaoteng Ma, Chenghao Li, Zewu Zheng, Qiyuan Zhang, Gao Huang, Jun Yang, and Qianchuan Zhao · 2021
Closest in time.