Fetching the paper…
Reading the bibliography…
This work focuses on off-policy evaluation (OPE) with function approximation in infinite-horizon undiscounted Markov decision processes (MDPs).
Large-scale markov decision problems via the linear programming dual
Yasin Abbasi-Yadkori, Peter L. Bartlett, Xi Chen, and Alan Malek · 1901
Earlier work this paper cites.
AlgaeDICE: Policy Gradient from Arbitrary Experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 1912
Earlier work this paper cites.
The rotation of eigenvectors by a perturbation. iii
Chandler Davis and William Morton Kahan · 1970
Earlier work this paper cites.
Simulation and the Monte Carlo method
RY Rubinstein · 1981
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Steven J Bradtke and Andrew G Barto · 1996
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Richard S Sutton · 1996
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Maximum entropy discrimination
Tommi Jaakkola, Marina Meila, and Tony Jebara · 2000
Earlier work this paper cites.
Dirichlet distribution
M Hazewinkel · 2001
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Cited alongside, same era.
Convergence results for some temporal difference methods based on least squares
Huizhen Yu and Dimitri P Bertsekas · 2009
Cited alongside, same era.
Value function approximation in reinforcement learning using the Fourier basis
George Konidaris, Sarah Osentoski, and Philip Thomas · 2011
Cited alongside, same era.
User-friendly tail bounds for sums of random matrices
Joel A Tropp · 2012
Cited alongside, same era.
Policy evaluation with temporal differences: A survey and comparison
Christoph Dann, Gerhard Neumann, Jan Peters, et al · 2014
Cited alongside, same era.
Off-policy learning with eligibility traces: A survey
Online linear quadratic control
Alon Cohen, Avinatan Hasidim, Tomer Koren, Nevena Lazić, Yishay Mansour, and Kunal Talwar · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Later among the works it cites.
On the sample complexity of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2019
Later among the works it cites.
A Kernel Loss for Solving the Bellman Equation
Yihao Feng, Lihong Li, and Qiang Liu · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Later among the works it cites.
Large-scale Markov Decision Processes with changing rewards
Adrian Rivera Cardoso, He Wang, and Huan Xu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthieu Geist and Bruno Scherrer · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Weighted importance sampling for off-policy learning with linear function approximation
A Rupam Mahmood, Hado P van Hasselt, and Richard S Sutton · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
A useful variant of the Davis–Kahan theorem for statisticians
Yi Yu, Tengyao Wang, and Richard J Samworth · 2015
Cited alongside, same era.
OpenAI Gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
CVXPY: A Python-embedded modeling language for convex optimization
Steven Diamond and Stephen Boyd · 2016
Cited alongside, same era.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin F Yang and Mengdi Wang · 2019
Later among the works it cites.
Minimax-optimal off-policy evaluation with linear function approximation
Yaqi Duan and Mengdi Wang · 2020
Closest in time.
Reinforcement learning via fenchel-rockafellar duality
Ofir Nachum and Bo Dai · 2020
Closest in time.
Batch stationary distribution estimation
Junfeng Wen, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Closest in time.
Q* Approximation Schemes for Batch Reinforcement Learning: A Theoretical Comparison, 2020
Tengyang Xie and Nan Jiang · 2020
Closest in time.
GenDICE: Generalized offline estimation of stationary values
Ruiyi Zhang, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Closest in time.