Off-policy policy gradient with state distribution correction
Original
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 1904
Earlier work this paper cites.
AlgaeDICE: Policy gradient from arbitrary experience
Original
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 1912
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Empirical Processes in M-estimation , volume 6
Sara Van de Geer · 2000
Earlier work this paper cites.
The probabilistic method
Jiří Matoušek and Jan Vondrák · 2001
Earlier work this paper cites.
GradientDICE: Rethinking generalized offline estimation of stationary values
Original
Shantong Zhang, Bo Liu, and Shimon Whiteson · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Finite time bounds for sampling based fitted value iteration
Csaba Szepesvári and Rémi Munos · 2005
Earlier work this paper cites.
Fitted Q-iteration in continuous action-space mdps
Andras Antos, Rémi Munos, and Csaba Szepesvari · 2007
Earlier work this paper cites.
Performance bounds in ℓ p \ell_{p} -norm for approximate value iteration
Rémi Munos · 2007
Earlier work this paper cites.
Variational policy gradient method for reinforcement learning with general utilities
Original
Junyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvari, and Mengdi Wang · 2007
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Variational analysis , volume 317
R Tyrrell Rockafellar and Roger J-B Wets · 2009
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Amir Massoud Farahmand, Rémi Munos, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
XuanLong Nguyen, Martin J Wainwright, and Michael I Jordan · 2010
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman · 2014
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Original
Stephane Ross and J Andrew Bagnell · 2014
Earlier work this paper cites.
Approximate policy iteration schemes: A comparison
Bruno Scherrer · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Batch learning from logged bandit feedback through counterfactual risk minimization
Adith Swaminathan and Thorsten Joachims · 2015
Earlier work this paper cites.
Learning from conditional distributions via dual embeddings
Bo Dai, Niao He, Yunpeng Pan, Byron Boots, and Le Song · 2017
Earlier work this paper cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisỳ, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Starcraft II: A new challenge for reinforcement learning
Original
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, Julian Schrittwieser, et al · 2017
Earlier work this paper cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Original
Jinglin Chen and Nan Jiang · 2019
Earlier work this paper cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S Du, Sham M Kakade, Ruosong Wang, and Lin F Yang · 2019
Earlier work this paper cites.
A kernel loss for solving the Bellman equation
Original
Yihao Feng, Lihong Li, and Qiang Liu · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.