Fetching the paper…
Reading the bibliography…
We study the problem of off-policy policy optimization in Markov decision processes, and develop a novel off-policy policy gradient method.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S Sutton, and Satinder P Singh · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Clinical data based optimal sti strategies for hiv: a reinforcement learning approach
Damien Ernst, Guy-Bart Stan, Jorge Goncalves, and Louis Wehenkel · 2006
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Convergent temporal-difference learning with arbitrary smooth function approximation
Shalabh Bhatnagar, Doina Precup, David Silver, Richard S Sutton, Hamid R Maei, and Csaba Szepesvári · 2009
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint , volume 48
Vivek S Borkar · 2009
Earlier work this paper cites.
Gradient Temporal-difference Learning Algorithms
Hamid Reza Maei · 2011
Earlier work this paper cites.
Thomas Degris, Martha White, and Richard S Sutton · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Sample efficient actor-critic with experience replay
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2016
Later among the works it cites.
Consistent on-line off-policy evaluation
Assaf Hallak and Shie Mannor · 2017
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudik · 2017
Later among the works it cites.
Sbeed: Convergent reinforcement learning with nonlinear function approximation
Bo Dai, Albert Shaw, Lihong Li, Lin Xiao, Niao He, Zhen Liu, Jianshu Chen, and Le Song · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Philip Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Cited alongside, same era.
Deep reinforcement learning through policy optimization
Pieter Abbeel and John Schulman · 2016
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Shixiang Shane Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard E. Turner, and Sergey Levine
Cited in the paper.
Interpolated policy gradient: Merging on-policy and off-policy gradient estimation for deep reinforcement learning
Shixiang Shane Gu, Timothy Lillicrap, Richard E Turner, Zoubin Ghahramani, Bernhard Schölkopf, and Sergey Levine
Cited in the paper.
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al · 2018
Later among the works it cites.
An off-policy policy gradient theorem using emphatic weightings
Ehsan Imani, Eric Graves, and Martha White · 2018
Later among the works it cites.
Policy optimization via importance sampling
Alberto Maria Metelli, Matteo Papini, Francesco Faccio, and Marcello Restelli · 2018
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Carles Gelada and Marc G Bellemare · 2019
Closest in time.
Generalized off-policy actor-critic
Shangtong Zhang, Wendelin Boehmer, and Shimon Whiteson · 2019
Closest in time.