Fetching the paper…
Reading the bibliography…
This paper studies the statistical theory of batch data reinforcement learning with function approximation.
On tail probabilities for martingales
Freedman, D. A · 1975
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Bertsekas, D. P., Bertsekas, D. P., Bertsekas, D. P., and Bertsekas, D. P · 1995
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D · 2000
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R · 2003
Earlier work this paper cites.
Bias and variance in value function estimation
Mannor, S., Simester, D., Sun, P., and Tsitsiklis, J. N · 2004
Earlier work this paper cites.
Model-based function approximation in reinforcement learning
Jong, N. K. and Stone, P · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V., Hayes, T. P., and Kakade, S. M · 2008
Earlier work this paper cites.
Freedman’s inequality for matrix martingales
Tropp, J. et al · 2011
Cited alongside, same era.
Modelling transition dynamics in mdps with rkhs embeddings
Grunewalder, S., Lever, G., Baldassarre, L., Pontil, M., and Gretton, A · 2012
Cited alongside, same era.
Batch mode reinforcement learning based on the synthesis of artificial trajectories
Fonteneau, R., Murphy, S. A., Wehenkel, L., and Ernst, D · 2013
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E · 2016
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., and Zhou, D · 2018
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
Dann, C., Li, L., Wei, W., and Brunskill, E · 2019
Later among the works it cites.
Off-policy policy gradient with state distribution correction
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E · 2019
Later among the works it cites.
DualDICE: Behavior-agnostic estimation of discounted stationary distribution corrections
Nachum, O., Chow, Y., Dai, B., and Li, L · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T., Ma, Y., and Wang, Y.-X · 2019
Later among the works it cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. F. and Wang, M · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Yin, M. and Wang, Y.-X · 2020
Closest in time.