Fetching the paper…
Reading the bibliography…
Actor-critic methods are widely used in offline reinforcement learning practice, but are not so well-understood theoretically.
Alternative theoretical frameworks for finite horizon discrete-time stochastic optimal control
Steven E Shreve and Dimitri P Bertsekas · 1978
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Dynamic programming and stochastic control
D. P. Bertsekas · 1995
Earlier work this paper cites.
Neuro-dynamic programming
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, Yishay Mansour, et al · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade et al · 2003
Earlier work this paper cites.
On actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Rémi Munos · 2003
Earlier work this paper cites.
Least squares policy evaluation algorithms with linear function approximation
A Nedić and Dimitri P Bertsekas · 2003
Earlier work this paper cites.
Error bounds for approximate value iteration
Rémi Munos · 2005
Earlier work this paper cites.
Fitted q-iteration in continuous action-space mdps
András Antos, Rémi Munos, and Csaba Szepesvári · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Self-normalized processes
V. H. de la Pena, T. L. Lai, and Q. M. Shao · 2009
Earlier work this paper cites.
Introduction to Nonparametric Estimation
A. B. Tsybakov · 2009
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Amir-massoud Farahmand, Csaba Szepesvári, and Rémi Munos · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, David Pal, and Csaba Szepesvari · 2011
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2014
Earlier work this paper cites.
Offline policy evaluation across representations with applications to educational games
Travis Mandel, Yun-En Liu, Sergey Levine, Emma Brunskill, and Zoran Popovic · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients, 2015
Nicolas Heess, Greg Wayne, David Silver, Timothy Lillicrap, Yuval Tassa, and Tom Erez · 2015
Earlier work this paper cites.
The information geometry of mirror descent
Garvesh Raskutti and Sayan Mukherjee · 2015
Earlier work this paper cites.
High confidence policy improvement
Philip Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Earlier work this paper cites.
Regularized policy iteration with nonparametric function spaces
Amir-massoud Farahmand, Mohammad Ghavamzadeh, Csaba Szepesvári, and Shie Mannor · 2016
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Earlier work this paper cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E. Schapire · 2017
Earlier work this paper cites.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Representation balancing mdps for off-policy policy evaluation
Yao Liu, Omer Gottesman, Aniruddh Raghu, Matthieu Komorowski, Aldo A Faisal, Finale Doshi-Velez, and Emma Brunskill · 2018
Earlier work this paper cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Cited alongside, same era.
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Cited alongside, same era.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine · 2019
Cited alongside, same era.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2020
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Noah Y Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
Masatoshi Uehara, Jiawei Huang, and Nan Jiang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nathan Kallus and Masatoshi Uehara · 2019
Cited alongside, same era.
Safe policy improvement with baseline bootstrapping
Romain Laroche, Paul Trichelair, and Remi Tachet Des Combes · 2019
Cited alongside, same era.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Cited alongside, same era.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Cited alongside, same era.
Preventing undesirable behavior of intelligent machines
Philip S Thomas, Bruno Castro da Silva, Andrew G Barto, Stephen Giguere, Yuriy Brun, and Emma Brunskill · 2019
Cited alongside, same era.
Doubly robust bias reduction in infinite horizon off-policy estimation
Ziyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou, and Qiang Liu · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint
Martin J Wainwright · 2019
Cited alongside, same era.
Gellert Weisz, Philip Amortila, and Csaba Szepesvári · 2020
Later among the works it cites.
Ziyu Wang, Alexander Novikov, Konrad Żołna, Jost Tobias Springenberg, Scott Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Heess, et al · 2020
Later among the works it cites.
Batch value-function approximation with only realizability
Tengyang Xie and Nan Jiang · 2020
Later among the works it cites.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Tengyang Xie and Nan Jiang · 2020
Later among the works it cites.
Near optimal provable uniform convergence in off-policy evaluation for reinforcement learning
Ming Yin, Yu Bai, and Yu-Xiang Wang · 2020
Later among the works it cites.
Off-policy evaluation via the regularized lagrangian
Mengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Lin F Yang and Mengdi Wang · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Ming Yin and Yu-Xiang Wang · 2020
Later among the works it cites.
Andrea Zanette · 2020
Later among the works it cites.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Matteo Pirotta, and Alessandro Lazaric · 2020
Later among the works it cites.
Gendice: Generalized offline estimation of stationary values
Ruiyi Zhang, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Later among the works it cites.
Variational policy gradient method for reinforcement learning with general utilities
Junyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvari, and Mengdi Wang · 2020
Later among the works it cites.
Learning near optimal policies with low inherent bellman error
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
Provably efficient reward-agnostic navigation with linear value iteration
Andrea Zanette, Alessandro Lazaric, Mykel J Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
Risk bounds and rademacher complexity in batch reinforcement learning
Yaqi Duan, Chi Jin, and Zhiyuan Li · 2021
Closest in time.
Bilinear classes: A structural framework for provable generalization in rl
Simon S Du, Sham M Kakade, Jason D Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Closest in time.
Bootstrapping statistical inference for off-policy evaluation
Botao Hao, Xiang Ji, Yaqi Duan, Hao Lu, Csaba Szepesvári, and Mengdi Wang · 2021
Closest in time.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Closest in time.
On the linear convergence of natural policy gradient algorithm
Sajad Khodadadian, Prakirt Raj Jhunjhunwala, Sushil Mahavir Varma, and Siva Theja Maguluri · 2021
Closest in time.
Guanghui Lan · 2021
Closest in time.
Model-free representation learning and exploration in low-rank mdps
Aditya Modi, Jinglin Chen, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2021
Closest in time.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Closest in time.
Pessimistic model-based offline rl: Pac bounds and posterior sampling under partial coverage, 2021
Masatoshi Uehara and Wen Sun · 2021
Closest in time.
Minimax model learning
Cameron Voloshin, Nan Jiang, and Yisong Yue · 2021
Closest in time.
Uncertainty weighted actor-critic for offline reinforcement learning
Yue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua Susskind, Jian Zhang, Ruslan Salakhutdinov, and Hanlin Goh · 2021
Closest in time.
Bellman-consistent pessimism for offline reinforcement learning, 2021
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal · 2021
Closest in time.