Fetching the paper…
Reading the bibliography…
In offline reinforcement learning (RL) we have no opportunity to explore so we must make assumptions that the data is sufficient to guide picking a good policy, taking the form of assuming some coverage, realizability, Bellman completeness, and/or hard margin (gap).
Fast learning rates for plug-in classifiers
Jean-Yves Audibert and Alexandre B Tsybakov · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Convex optimization theory , volume 1
Dimitri Bertsekas · 2009
Earlier work this paper cites.
Generalizing evidence from randomized clinical trials to target populations: the actg 320 trial
Stephen R Cole and Elizabeth A Stuart · 2010
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
Brian D Ziebart, J Andrew Bagnell, and Anind K Dey · 2010
Earlier work this paper cites.
The multi-armed bandit problem with covariates
Vianney Perchet and Philippe Rigollet · 2013
Earlier work this paper cites.
External validity: From do-calculus to transportability across populations
Judea Pearl and Elias Bareinboim · 2014
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2015
Earlier work this paper cites.
On the properties of the softmax function with application in game theory and reinforcement learning
Bolin Gao and Lacra Pavel · 2017
Earlier work this paper cites.
A unified view of entropy-regularized markov decision processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Earlier work this paper cites.
Equivalence between policy gradients and soft q-learning
John Schulman, Xi Chen, and Pieter Abbeel · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Earlier work this paper cites.
Generalizing causal inferences from individuals in randomized trials to all trial-eligible individuals
Issa J Dahabreh, Sarah E Robertson, Eric J Tchetgen, Elizabeth A Stuart, and Miguel A Hernán · 2019
Earlier work this paper cites.
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Earlier work this paper cites.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Max Simchowitz and Kevin G Jamieson · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
Minimax value interval for off-policy evaluation and policy optimization
Nan Jiang and Jiawei Huang · 2020
Cited alongside, same era.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
Bellman-consistent pessimism for offline reinforcement learning
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal · 2021
Later among the works it cites.
Towards instance-optimal offline reinforcement learning with pessimism
Ming Yin and Yu-Xiang Wang · 2021
Later among the works it cites.
Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning
Chenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhihong Deng, Animesh Garg, Peng Liu, and Zhaoran Wang · 2022
Later among the works it cites.
Offline reinforcement learning under value and density-ratio realizability: the power of gaps
Jinglin Chen and Nan Jiang · 2022
Later among the works it cites.
Xiaohong Chen and Zhengling Qi · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Performance guarantees for policy learning
Alex Luedtke and Antoine Chambaz · 2020
Cited alongside, same era.
Reinforcement learning via fenchel-rockafellar duality
Ofir Nachum and Bo Dai · 2020
Cited alongside, same era.
Off-policy evaluation and learning for external validity under a covariate shift
Masatoshi Uehara, Masahiro Kato, and Shota Yasui · 2020
Cited alongside, same era.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Tengyang Xie and Nan Jiang · 2020
Cited alongside, same era.
Off-policy evaluation via the regularized lagrangian
Mengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Cited alongside, same era.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Cited alongside, same era.
Mitigating covariate shift in imitation learning via offline data with partial coverage
Jonathan Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi, and Wen Sun · 2021
Cited alongside, same era.
Later among the works it cites.
Fast rates for contextual linear optimization
Yichun Hu, Nathan Kallus, and Xiaojie Mao · 2022
Later among the works it cites.
Beyond the return: Off-policy function estimation under user-specified error-measuring distributions
Audrey Huang and Nan Jiang · 2022
Later among the works it cites.
Pessimism for offline linear contextual bandits using lp confidence sets
Gene Li, Cong Ma, and Nathan Srebro · 2022
Later among the works it cites.
On instance-dependent bounds for offline reinforcement learning with linear function approximation
Thanh Nguyen-Tang, Ming Yin, Sunil Gupta, Svetha Venkatesh, and Raman Arora · 2022
Later among the works it cites.
Optimal conservative offline rl with general function approximation via augmented lagrangian
Paria Rashidinejad, Hanlin Zhu, Kunhe Yang, Stuart Russell, and Jiantao Jiao · 2022
Later among the works it cites.
On gap-dependent bounds for offline reinforcement learning
Xinqi Wang, Qiwen Cui, and Simon S Du · 2022
Later among the works it cites.
Gap-dependent unsupervised exploration for reinforcement learning
Jingfeng Wu, Vladimir Braverman, and Lin Yang · 2022
Later among the works it cites.
Bellman residual orthogonalization for offline reinforcement learning
Andrea Zanette and Martin J Wainwright · 2022
Later among the works it cites.
Offline reinforcement learning with realizability and single-policy concentrability
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, and Jason Lee · 2022
Later among the works it cites.
Corruption-robust offline reinforcement learning
Xuezhou Zhang, Yiding Chen, Xiaojin Zhu, and Wen Sun · 2022
Later among the works it cites.
Revisiting the linear-programming framework for offline rl with general function approximation
Asuman E Ozdaglar, Sarath Pattathil, Jiawei Zhang, and Kaiqing Zhang · 2023
Closest in time.
Importance weighted actor-critic for optimal conservative offline reinforcement learning
Hanlin Zhu, Paria Rashidinejad, and Jiantao Jiao · 2023
Closest in time.