Fetching the paper…
Reading the bibliography…
In reinforcement learning (RL), offline learning decoupled learning from data collection and is useful in dealing with exploration-exploitation tradeoff and enables data reuse in many applications.
Discrete dynamic programming
David Blackwell · 1962
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Dynamic Programming and Optimal Control
D.P. Bertsekas · 1995
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
The ode method for convergence of stochastic approximation and reinforcement learning
Vivek S Borkar and Sean P Meyn · 2000
Earlier work this paper cites.
Call admission control and routing in integrated services networks using neuro-dynamic programming
Peter Marbach, Oliver Mihatsch, and John N Tsitsiklis · 2000
Earlier work this paper cites.
The necessity of average rewards in cooperative multirobot learning
Poj Tangamchit, John M Dolan, and Pradeep K Khosla · 2002
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications
Harold Kushner and G George Yin · 2003
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Earlier work this paper cites.
Approximate policy iteration: A survey and some new methods
Dimitri P Bertsekas · 2011
Earlier work this paper cites.
The generalization ability of online algorithms for dependent data
Alekh Agarwal and John C Duchi · 2012
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Markov chains and mixing times
David A Levin and Yuval Peres · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Jalaj Bhandari, Daniel Russo, and Raghav Singal · 2018
Finite-time error bounds for linear stochastic approximation andtd learning
Rayadurgam Srikant and Lei Ying · 2019
Later among the works it cites.
Two time-scale off-policy td learning: Non-asymptotic analysis over markovian samples
Tengyu Xu, Shaofeng Zou, and Yingbin Liang · 2019
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Finite-sample analysis of proximal gradient td algorithms
Bo Liu, Ji Liu, Mohammad Ghavamzadeh, Sridhar Mahadevan, and Marek Petrik · 2020
Later among the works it cites.
Black-box off-policy estimation for infinite-horizon reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Finite sample analysis of two-timescale stochastic approximation with applications to reinforcement learning
Gal Dalal, Gugan Thoppe, Balázs Szörényi, and Shie Mannor · 2018
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Cited alongside, same era.
On markov chain gradient descent
Tao Sun, Yuejiao Sun, and Wotao Yin · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
R. Sutton and A. Barto · 2018
Cited alongside, same era.
Finite sample analysis of the gtd policy evaluation algorithms in markov setting
Yue Wang, Wei Chen, Yuting Liu, Zhi-Ming Ma, and Tie-Yan Liu · 2018
Cited alongside, same era.
Finite-time performance bounds and adaptive learning rate selection for two time-scale reinforcement learning
Harsh Gupta, R Srikant, and Lei Ying · 2019
Cited alongside, same era.
Ali Mousavi, Lihong Li, Qiang Liu, and Denny Zhou · 2020
Later among the works it cites.
Finite-time analysis of asynchronous stochastic approximation and q q -learning
Guannan Qu and Adam Wierman · 2020
Later among the works it cites.
Gradientdice: Rethinking generalized offline estimation of stationary values
Shangtong Zhang, Bo Liu, and Shimon Whiteson · 2020
Later among the works it cites.
A lyapunov theory for finite-sample guarantees of asynchronous q-learning and td-learning variants
Zaiwei Chen, Siva Theja Maguluri, Sanjay Shakkottai, and Karthikeyan Shanmugam · 2021
Later among the works it cites.
Learning and planning in average-reward markov decision processes
Yi Wan, Abhishek Naik, and Richard S Sutton · 2021
Later among the works it cites.
Near-optimal provable uniform convergence in offline policy evaluation for reinforcement learning
Ming Yin, Yu Bai, and Yu-Xiang Wang · 2021
Later among the works it cites.
Average-reward off-policy policy evaluation with function approximation
Shangtong Zhang, Yi Wan, Richard S Sutton, and Shimon Whiteson · 2021
Later among the works it cites.