Fetching the paper…
Reading the bibliography…
We study the off-policy evaluation (OPE) problem in reinforcement learning with linear function approximation, which aims to estimate the value function of a target policy based on the offline data collected by a behavior policy.
Algaedice: Policy gradient from arbitrary experience
O. Nachum, B. Dai, I. Kostrikov, Y. Chow, L. Li, and D. Schuurmans · 1912
Earlier work this paper cites.
On tail probabilities for martingales
D. A. Freedman · 1975
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup · 2000
Earlier work this paper cites.
A natural policy gradient
S. M. Kakade · 2001
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
D. Precup, R. S. Sutton, and S. Dasgupta · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
R. Vershynin · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
M. Dudík, J. Langford, and L. Li · 2011
Earlier work this paper cites.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
L. Li, W. Chu, J. Langford, and X. Wang · 2011
Earlier work this paper cites.
Testing different reinforcement learning configurations for financial trading: Introduction and applications
F. Bertoluzzo and M. Corazza · 2012
Earlier work this paper cites.
Is pessimism provably efficient for offline rl?
Y. Jin, Z. Yang, and Z. Wang · 2012
Earlier work this paper cites.
Batch reinforcement learning
S. Lange, T. Gabel, and M. Riedmiller · 2012
Earlier work this paper cites.
User-friendly tail bounds for sums of random matrices
J. A. Tropp · 2012
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
D. Charles, M. Chickering, and P. Simard · 2013
Earlier work this paper cites.
Safe policy iteration
M. Pirotta, M. Restelli, A. Pecorino, and D. Calandriello · 2013
Earlier work this paper cites.
Automatic ad format selection via contextual bandits
L. Tang, R. Rosales, A. Singh, and D. Agarwal · 2013
Cited alongside, same era.
Toward minimax off-policy value estimation
L. Li, R. Munos, and C. Szepesvári · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
P. Thomas and E. Brunskill · 2016
Cited alongside, same era.
Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing
P. S. Thomas, G. Theocharous, M. Ghavamzadeh, I. Durugkar, and E. Brunskill · 2017
Cited alongside, same era.
More robust doubly robust off-policy evaluation
M. Farajtabar, Y. Chow, and M. Ghavamzadeh · 2018
Coindice: Off-policy confidence interval estimation
B. Dai, O. Nachum, Y. Chow, L. Li, C. Szepesvári, and D. Schuurmans · 2020
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
S. S. Du, S. M. Kakade, R. Wang, and L. F. Yang · 2020
Later among the works it cites.
Minimax-optimal off-policy evaluation with linear function approximation
Y. Duan, Z. Jia, and M. Wang · 2020
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Z. Jia, L. Yang, C. Szepesvari, and M. Wang · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Q. Liu, L. Li, Z. Tang, and D. Zhou · 2018
Cited alongside, same era.
Deep reinforcement learning for vision-based robotic grasping: A simulated comparative evaluation of off-policy methods
D. Quillen, E. Jang, O. Nachum, C. Finn, J. Ibarz, and S. Levine · 2018
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2019
Cited alongside, same era.
N. Kallus and M. Uehara · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, G. Tucker, and S. Levine · 2019
Cited alongside, same era.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
T. Xie, Y. Ma, and Y.-X. Wang · 2019
Cited alongside, same era.
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
A unifying view of optimism in episodic reinforcement learning
G. Neu and C. Pike-Burke · 2020
Later among the works it cites.
Doubly robust bias reduction in infinite horizon off-policy estimation
Z. Tang, Y. Feng, L. Li, D. Zhou, and Q. Liu · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
M. Yin and Y.-X. Wang · 2020
Later among the works it cites.
Learning near optimal policies with low inherent bellman error
A. Zanette, A. Lazaric, M. Kochenderfer, and E. Brunskill · 2020
Later among the works it cites.
Gendice: Generalized offline estimation of stationary values
R. Zhang, B. Dai, L. Li, and D. Schuurmans · 2020
Later among the works it cites.
L. Chen, B. Scherrer, and P. L. Bartlett · 2021
Closest in time.
Logarithmic regret for reinforcement learning with linear function approximation
J. He, D. Zhou, and Q. Gu · 2021
Closest in time.
Learning stochastic shortest path with linear function approximation
Y. Min, J. He, T. Wang, and Q. Gu · 2021
Closest in time.
Z. Zhang, J. Yang, X. Ji, and S. S. Du · 2021
Closest in time.