Fetching the paper…
Reading the bibliography…
We propose A-Crab (Actor-Critic Regularized by Average Bellman error), a new practical algorithm for offline reinforcement learning (RL) in complex environments with insufficient data coverage.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 1904
Earlier work this paper cites.
AlgaeDICE: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 1912
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
The probabilistic method
Jiří Matoušek and Jan Vondrák · 2001
Earlier work this paper cites.
GradientDICE: Rethinking generalized offline estimation of stationary values
Shantong Zhang, Bo Liu, and Shimon Whiteson · 2001
Earlier work this paper cites.
Finite time bounds for sampling based fitted value iteration
Csaba Szepesvári and Rémi Munos · 2005
Earlier work this paper cites.
Fitted Q-iteration in continuous action-space mdps
Andras Antos, Rémi Munos, and Csaba Szepesvari · 2007
Earlier work this paper cites.
Performance bounds in ℓ p \ell_{p} -norm for approximate value iteration
Rémi Munos · 2007
Earlier work this paper cites.
Variational policy gradient method for reinforcement learning with general utilities
Junyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvari, and Mengdi Wang · 2007
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Online markov decision processes
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2009
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Amir Massoud Farahmand, Rémi Munos, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Approximate policy iteration schemes: A comparison
Bruno Scherrer · 2014
Earlier work this paper cites.
Batch learning from logged bandit feedback through counterfactual risk minimization
Adith Swaminathan and Thorsten Joachims · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Earlier work this paper cites.
A kernel loss for solving the Bellman equation
Yihao Feng, Lihong Li, and Qiang Liu · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard · 2019
Earlier work this paper cites.
On value functions and the agent-environment boundary
Nan Jiang · 2019
Earlier work this paper cites.
Stabilizing off-policy Q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Earlier work this paper cites.
Deepee: Joint optimization of job scheduling and cooling control for data center energy efficiency using deep reinforcement learning
Yongyi Ran, Han Hu, Xin Zhou, and Yonggang Wen · 2019
Earlier work this paper cites.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Cited alongside, same era.
D4RL: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
EMaQ: Expected-max Q-learning operator for simple yet effective offline and online RL
Seyed Kamyar Seyed Ghasemipour, Dale Schuurmans, and Shixiang Shane Gu · 2020
Cited alongside, same era.
Bellman-consistent pessimism for offline reinforcement learning
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal · 2021
Later among the works it cites.
Towards instance-optimal offline reinforcement learning with pessimism
Ming Yin and Yu-Xiang Wang · 2021
Later among the works it cites.
COMBO: Conservative offline model-based policy optimization
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Andrea Zanette, Martin J Wainwright, and Emma Brunskill · 2021
Later among the works it cites.
Offline reinforcement learning under value and density-ratio realizability: The power of gaps
Jinglin Chen and Nan Jiang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Cited alongside, same era.
Conservative Q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Batch policy learning in average reward Markov decision processes
Peng Liao, Zhengling Qi, and Susan Murphy · 2020
Cited alongside, same era.
Provably good batch reinforcement learning without great exploration
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2020
Cited alongside, same era.
Chip placement with deep reinforcement learning
Azalia Mirhoseini, Anna Goldie, Mustafa Yazgan, Joe Jiang, Ebrahim Songhori, Shen Wang, Young-Joon Lee, Eric Johnson, Omkar Pathak, Sungmin Bae, et al · 2020
Cited alongside, same era.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2020
Cited alongside, same era.
Discovering reinforcement learning algorithms
Junhyuk Oh, Matteo Hessel, Wojciech M Czarnecki, Zhongwen Xu, Hado P van Hasselt, Satinder Singh, and David Silver · 2020
Cited alongside, same era.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Noah Y Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, and Martin Riedmiller · 2020
Cited alongside, same era.
Adversarially trained actor critic for offline reinforcement learning
Ching-An Cheng, Tengyang Xie, Nan Jiang, and Alekh Agarwal · 2022
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de Las Casas, et al · 2022
Later among the works it cites.
Discovering faster matrix multiplication algorithms with reinforcement learning
Alhussein Fawzi, Matej Balog, Aja Huang, Thomas Hubert, Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Francisco J R Ruiz, Julian Schrittwieser, Grzegorz Swirszcz, et al · 2022
Later among the works it cites.
Model-based offline reinforcement learning with pessimism-modulated dynamics belief
Kaiyang Guo, Yunfeng Shao, and Yanhui Geng · 2022
Later among the works it cites.
Revisiting the linear-programming framework for offline RL with general function approximation
Asuman Ozdaglar, Sarath Pattathil, Jiawei Zhang, and Kaiqing Zhang · 2022
Later among the works it cites.
Optimal conservative offline RL with general function approximation via augmented Lagrangian
Paria Rashidinejad, Hanlin Zhu, Kunhe Yang, Stuart Russell, and Jiantao Jiao · 2022
Later among the works it cites.
Offline reinforcement learning as anti-exploration
Shideh Rezaeifar, Robert Dadashi, Nino Vieillard, Léonard Hussenot, Olivier Bachem, Olivier Pietquin, and Matthieu Geist · 2022
Later among the works it cites.
RAMBO-RL: Robust adversarial model-based offline reinforcement learning
Marc Rigter, Bruno Lacerda, and Nick Hawes · 2022
Later among the works it cites.
Tackling climate change with machine learning
David Rolnick, Priya L Donti, Lynn H Kaack, Kelly Kochanski, Alexandre Lacoste, Kris Sankaran, Andrew Slavin Ross, Nikola Milojevic-Dupont, Natasha Jaques, Anna Waldman-Brown, et al · 2022
Later among the works it cites.
Laixi Shi and Yuejie Chi · 2022
Later among the works it cites.
S4RL: Surprisingly simple self-supervision for offline reinforcement learning in robotics
Samarth Sinha, Ajay Mandlekar, and Animesh Garg · 2022
Later among the works it cites.
Hybrid RL: Using both offline and online data can make RL efficient
Yuda Song, Yifei Zhou, Ayush Sekhari, J Andrew Bagnell, Akshay Krishnamurthy, and Wen Sun · 2022
Later among the works it cites.
Leveraging factored action spaces for efficient offline reinforcement learning in healthcare
Shengpu Tang, Maggie Makar, Michael Sjoding, Finale Doshi-Velez, and Jenna Wiens · 2022
Later among the works it cites.
On gap-dependent bounds for offline reinforcement learning
Xinqi Wang, Qiwen Cui, and Simon S Du · 2022
Later among the works it cites.
Armor: A model-based framework for improving arbitrary baseline policies with offline data
Tengyang Xie, Mohak Bhardwaj, Nan Jiang, and Ching-An Cheng · 2022
Later among the works it cites.
The efficacy of pessimism in asynchronous Q-learning
Yuling Yan, Gen Li, Yuxin Chen, and Jianqing Fan · 2022
Later among the works it cites.
Ming Yin, Yaqi Duan, Mengdi Wang, and Yu-Xiang Wang · 2022
Later among the works it cites.
Offline reinforcement learning with realizability and single-policy concentrability
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, and Jason Lee · 2022
Later among the works it cites.
Corruption-robust offline reinforcement learning
Xuezhou Zhang, Yiding Chen, Xiaojin Zhu, and Wen Sun · 2022
Later among the works it cites.
A primal-dual-critic algorithm for offline constrained reinforcement learning
Kihyuk Hong, Yuhang Li, and Ambuj Tewari · 2023
Closest in time.