Fetching the paper…
Reading the bibliography…
In offline reinforcement learning (RL), we seek to utilize offline data to evaluate (or learn) policies in scenarios where the data are collected from a distribution that substantially differs from that of the target policy to be evaluated.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
O. Nachum, Y. Chow, B. Dai, and L. Li · 1906
Earlier work this paper cites.
Algaedice: Policy gradient from arbitrary experience
O. Nachum, B. Dai, I. Kostrikov, Y. Chow, L. Li, and D. Schuurmans · 1912
Earlier work this paper cites.
Neuro-dynamic programming: an overview
D. P. Bertsekas and J. N. Tsitsiklis · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
G. J. Gordon · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. J. Bradtke and A. G. Barto · 1996
Earlier work this paper cites.
Stable fitted reinforcement learning
G. J. Gordon · 1996
Earlier work this paper cites.
Approximate solutions to markov decision processes
G. J. Gordon · 1999
Earlier work this paper cites.
Barycentric interpolators for continuous space and time reinforcement learning
R. Munos and A. W. Moore · 1999
Earlier work this paper cites.
Kernel-based reinforcement learning
D. Ormoneit and Ś. Sen · 2002
Earlier work this paper cites.
Gendice: Generalized offline estimation of stationary values
R. Zhang, B. Dai, L. Li, and D. Schuurmans · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
R. Munos · 2003
Earlier work this paper cites.
Condition numbers of gaussian random matrices
Z. Chen and J. J. Dongarra · 2005
Earlier work this paper cites.
Finite time bounds for sampling based fitted value iteration
C. Szepesvári and R. Munos · 2005
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi, B. Recht, et al · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
A. Antos, C. Szepesvári, and R. Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
R. Munos and C. Szepesvári · 2008
Earlier work this paper cites.
Doubly robust policy evaluation and learning
M. Dudík, J. Langford, and L. Li · 2011
Earlier work this paper cites.
A tail inequality for quadratic forms of subgaussian random vectors
D. Hsu, S. Kakade, T. Zhang, et al · 2012
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
S. Ross and D. Bagnell · 2012
Earlier work this paper cites.
Offline policy evaluation across representations with applications to educational games
T. Mandel, Y.-E. Liu, S. Levine, E. Brunskill, and Z. Popovic · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson · 2014
Earlier work this paper cites.
Toward minimax off-policy value estimation
L. Li, R. Munos, and C. Szepesvari · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
High-confidence off-policy evaluation
P. S. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Earlier work this paper cites.
An introduction to matrix concentration inequalities
J. A. Tropp · 2015
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
P. Thomas and E. Brunskill · 2016
Cited alongside, same era.
Minimax weight and q-function learning for off-policy evaluation
M. Uehara and N. Jiang · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
T. Xie, Y. Ma, and Y.-X. Wang · 2019
Later among the works it cites.
Deep inverse reinforcement learning for sepsis treatment
C. Yu, G. Ren, and J. Liu · 2019
Later among the works it cites.
Limiting extrapolation in linear approximate value iteration
A. Zanette, A. Lazaric, M. J. Kochenderfer, and E. Brunskill · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
R. Agarwal, D. Schuurmans, and M. Norouzi · 2020
Later among the works it cites.
A variant of the wang-foster-kakade lower bound for the discounted setting
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Guo, P. S. Thomas, and E. Brunskill · 2017
Cited alongside, same era.
Towards generalization and simplicity in continuous control
A. Rajeswaran, K. Lowrey, E. V. Todorov, and S. M. Kakade · 2017
Cited alongside, same era.
Boosted fitted q-iteration
S. Tosatto, M. Pirotta, C. d’Eramo, and M. Restelli · 2017
Cited alongside, same era.
Optimal and adaptive off-policy evaluation in contextual bandits
Y.-X. Wang, A. Agarwal, and M. Dudık · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
More robust doubly robust off-policy evaluation
M. Farajtabar, Y. Chow, and M. Ghavamzadeh · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Cited alongside, same era.
P. Amortila, N. Jiang, and T. Xie · 2020
Later among the works it cites.
Semantic visual navigation by watching youtube videos
M. Chang, A. Gupta, and S. Gupta · 2020
Later among the works it cites.
Minimax-optimal off-policy evaluation with linear function approximation
Y. Duan, Z. Jia, and M. Wang · 2020
Later among the works it cites.
Accountable off-policy evaluation with kernel bellman statistics
Y. Feng, T. Ren, Z. Tang, and Q. Liu · 2020
Later among the works it cites.
Way off-policy batch deep reinforcement learning of human preferences in dialog, 2020
N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2020
Later among the works it cites.
Minimax confidence interval for off-policy evaluation and policy optimization
N. Jiang and J. Huang · 2020
Later among the works it cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
N. Kallus and M. Uehara · 2020
Later among the works it cites.
Morel : Model-based offline reinforcement learning, 2020
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
Hyperparameter selection for offline reinforcement learning
T. L. Paine, C. Paduraru, A. Michi, C. Gulcehre, K. Zolna, A. Novikov, Z. Wang, and N. de Freitas · 2020
Later among the works it cites.
Offline reinforcement learning from images with latent space models
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn · 2020
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
N. Y. Siegel, J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, and M. Riedmiller · 2020
Later among the works it cites.
Behavior regularized offline reinforcement learning, 2020
Y. Wu, G. Tucker, and O. Nachum · 2020
Later among the works it cites.
Batch value-function approximation with only realizability
T. Xie and N. Jiang · 2020
Later among the works it cites.
Off-policy evaluation via the regularized lagrangian
M. Yang, O. Nachum, B. Dai, L. Li, and D. Schuurmans · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
A. Zanette · 2020
Later among the works it cites.
What are the statistical limits of offline rl with linear function approximation?
R. Wang, D. P. Foster, and S. M. Kakade · 2021
Closest in time.
Combo: Conservative offline model-based policy optimization
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn · 2021
Closest in time.