Fetching the paper…
Reading the bibliography…
Policy gradient methods in reinforcement learning update policy parameters by taking steps in the direction of an estimated gradient of policy value.
Orthogonal statistical learning
Foster, D. J. and V. Syrgkanis (2019) · 1901
Earlier work this paper cites.
Learning when-to-treat policies
Nie, X., E. Brunskill, and S. Wager (2019) · 1905
Earlier work this paper cites.
Global optimality guarantees for policy gradient methods
Bhandari, J. and D. Russo (2019) · 1906
Earlier work this paper cites.
Optimality and approximation with policy gradient methods in markov decision processes
Agarwal, A., S. Kakade, J. Lee, and G. Mahajan (2019) · 1908
Earlier work this paper cites.
Trajectory-wise control variates for variance reduction in policy gradient methods
Cheng, C.-A., X. Yan, and B. Boots (2019) · 1908
Earlier work this paper cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Kallus, N. and M. Uehara (2019a) · 1908
Earlier work this paper cites.
Kallus, N. and M. Uehara (2019b) · 1909
Earlier work this paper cites.
From importance sampling to doubly robust policy gradient
Huang, J. and N. Jiang (2019) · 1910
Earlier work this paper cites.
Algaedice: Policy gradient from arbitrary experience
Dai, B., I. Kostrikov, Y. Chow, L. Li, and D. Schuurmans (2019) · 1912
Earlier work this paper cites.
A characterization of limiting distributions of regular estimates
Hájek, J. (1970) · 1970
Earlier work this paper cites.
Consistent estimation of the influence function of locally asymptotically linear estimators
Klaassen, C. A. (1987) · 1987
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Efficient and adaptive estimation for semiparametric models
Bickel, P. J., C. A. Klaassen, Y. Ritov, and J. A. Wellner (1993) · 1993
Earlier work this paper cites.
Intra-option Learning about Temporally Abstract Actions
Sutton, R. S., D. Precup, and S. P. Singh (1998) · 1998
Earlier work this paper cites.
Asymptotic statistics
van der Vaart, A. W. (1998) · 1998
Earlier work this paper cites.
Eligibility Traces for Off-Policy Policy Evaluation
Precup, D., R. S. Sutton, and S. P. Singh (2000) · 2000
Earlier work this paper cites.
Infinite-Horizon Policy-Gradient Estimation
Baxter, J. and P. L. Bartlett (2001) · 2001
Earlier work this paper cites.
A nonparametric offpolicy policy gradient
Tosatto, S., J. Carvalho, H. Abdulsamad, and J. Peters (2020) · 2001
Earlier work this paper cites.
Unified Methods for Censored Longitudinal Data and Causality
van Der Laan, M. J. and J. M. Robins (2003) · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, E., P. L. Bartlett, and J. Baxter (2004) · 2004
Earlier work this paper cites.
Local rademacher complexities
Bartlett, P. L., O. Bousquet, and S. Mendelson (2005) · 2005
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Nesterov, Y. and B. Polyak (2006) · 2006
Cited alongside, same era.
Semiparametric Theory and Missing Data
Tsiatis, A. A. (2006) · 2006
Cited alongside, same era.
Chapter 76 large sample sieve estimation of semi-nonparametric models
Chen, X. (2007) · 2007
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., C. Szepesvári, and R. Munos (2008) · 2008
Cited alongside, same era.
Finite-time bounds for fitted value iteration
Munos, R. and C. Szepesvári (2008) · 2008
Cited alongside, same era.
Derivatives of logarithmic stationary distributions for policy gradient reinforcement learning
Morimura, T., E. Uchibe, J. Yoshimoto, J. Peters, and K. Doya (2010) · 2010
An off-policy policy gradient theorem using emphatic weightings
Imani, E., E. Graves, and M. White (2018) · 2018
Later among the works it cites.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Jiang, N. and A. Agarwal (2018) · 2018
Later among the works it cites.
Balanced policy evaluation and learning
Kallus, N. (2018) · 2018
Later among the works it cites.
Confounding-robust policy improvement
Kallus, N. and A. Zhou (2018) · 2018
Later among the works it cites.
Convergence guarantees for a class of non-convex and non-smooth optimization problems
Khamaru, K. and M. Wainwright (2018) · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., L. Li, Z. Tang, and D. Zhou (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On a connection between importance sampling and the likelihood ratio policy gradient
Tang, J. and P. Abbeel (2010) · 2010
Cited alongside, same era.
Off-policy actor-critic
Degris, T., M. White, and R. Sutton (2012) · 2012
Cited alongside, same era.
A survey on policy search for robotics
Deisenroth, M., G. Neumann, and J. Peters (2013) · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller (2014) · 2014
Cited alongside, same era.
Introduction to online convex optimization
Hazan, E. (2015) · 2015
Cited alongside, same era.
Gradient estimation using stochastic computation graphs
Schulman, J., N. Heess, T. Weber, and P. Abbeel (2015) · 2015
Cited alongside, same era.
Policy optimization via importance sampling
Metelli, A. M., M. Papini, F. Faccio, and M. Restelli (2018) · 2018
Later among the works it cites.
Stochastic variance-reduced policy gradient
Papini, M., D. Binaghi, G. Canonaco, M. Pirotta, and M. Restelli (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and A. G. Barto (2018) · 2018
Later among the works it cites.
Targeted Learning :Causal Inference for Observational and Experimental Data
van der Laan, M. J. and S. Rose (2018) · 2018
Later among the works it cites.
Variance reduction for policy gradient with action-dependent factorized baselines
Wu, C., A. Rajeswaran, Y. Duan, V. Kumar, A. Bayen, S. Kakade, I. Mordatch, and P. Abbeel (2018) · 2018
Later among the works it cites.
Offline multi-action policy learning: Generalization and optimization
Zhou, Z., S. Athey, and S. Wager (2018) · 2018
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J. and N. Jiang (2019) · 2019
Later among the works it cites.
Top-k off-policy correction for a reinforce recommender system
Chen, M., A. Beutel, P. Covington, S. Jain, F. Belletti, and E. Chi (2019) · 2019
Later among the works it cites.
Guidelines for reinforcement learning in healthcare
Gottesman, O., F. Johansson, M. Komorowski, A. Faisal, D. Sontag, F. Doshi-Velez, and L. A. Celi (2019) · 2019
Later among the works it cites.
Off-policy policy gradient with state distribution correction
Liu, Y., A. Swaminathan, A. Agarwal, and E. Brunskill (2019) · 2019
Later among the works it cites.
High-Dimensional Statistics : A Non-Asymptotic Viewpoint
Wainwright, M. J. (2019) · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T., Y. Ma, and Y.-X. Wang (2019) · 2019
Later among the works it cites.
A gentle introduction to empirical process theoryand applications
Sen, B. (2018) · 2020
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., D. Meger, and D. Precup (2019) · 2062
Closest in time.