Fetching the paper…
Reading the bibliography…
We propose and analyze a reinforcement learning principle that approximates the Bellman equations by enforcing their validity only along an user-defined space of test functions.
Series solution of some problems of elastic equilibrium of rods and plates
Boris Grigoryevich Galerkin · 1915
Earlier work this paper cites.
Computational Galerkin methods
Clive AJ Fletcher · 1984
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Dynamic programming and stochastic control
D. P. Bertsekas · 1995
Earlier work this paper cites.
Dynamic programming and stochastic control
D.P. Bertsekas · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Steven J Bradtke and Andrew G Barto · 1996
Earlier work this paper cites.
Neuro-dynamic programming
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade et al · 2003
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Rémi Munos · 2003
Earlier work this paper cites.
Error bounds for approximate value iteration
Rémi Munos · 2005
Earlier work this paper cites.
Fitted Q-iteration in continuous action-space MDPs
András Antos, Rémi Munos, and Csaba Szepesvári · 2007
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Partial differential equations
Lawrence C Evans · 2010
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Amir-massoud Farahmand, Csaba Szepesvári, and Rémi Munos · 2010
Earlier work this paper cites.
Error bounds for approximations from projected linear equations
H. Yu and D. P. Bertsekas · 2010
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
Regularized policy iteration with nonparametric function spaces
Amir-massoud Farahmand, Mohammad Ghavamzadeh, Csaba Szepesvári, and Shie Mannor · 2016
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Earlier work this paper cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Earlier work this paper cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E. Schapire · 2017
Earlier work this paper cites.
One hundred years of the Galerkin method
Sergey Repin · 2017
Earlier work this paper cites.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Earlier work this paper cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Cited alongside, same era.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine · 2019
Cited alongside, same era.
Nathan Kallus and Masatoshi Uehara · 2019
Near optimal provable uniform convergence in off-policy evaluation for reinforcement learning
Ming Yin, Yu Bai, and Yu-Xiang Wang · 2020
Later among the works it cites.
Off-policy evaluation via the regularized lagrangian
Mengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Ming Yin and Yu-Xiang Wang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Safe policy improvement with baseline bootstrapping
Romain Laroche, Paul Trichelair, and Remi Tachet Des Combes · 2019
Cited alongside, same era.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Cited alongside, same era.
Doubly robust bias reduction in infinite horizon off-policy estimation
Ziyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou, and Qiang Liu · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint
Martin J Wainwright · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Cited alongside, same era.
Andrea Zanette · 2020
Later among the works it cites.
Gendice: Generalized offline estimation of stationary values
Ruiyi Zhang, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Later among the works it cites.
Gradientdice: Rethinking generalized offline estimation of stationary values
Shangtong Zhang, Bo Liu, and Shimon Whiteson · 2020
Later among the works it cites.
Risk bounds and rademacher complexity in batch reinforcement learning
Yaqi Duan, Chi Jin, and Zhiyuan Li · 2021
Later among the works it cites.
Offline reinforcement learning: Fundamental barriers for value function approximation, 2021
Dylan J. Foster, Akshay Krishnamurthy, David Simchi-Levi, and Yunzong Xu · 2021
Later among the works it cites.
Introduction to online convex optimization, 2021
Elad Hazan · 2021
Later among the works it cites.
Bootstrapping statistical inference for off-policy evaluation
Botao Hao, Xiang Ji, Yaqi Duan, Hao Lu, Csaba Szepesvári, and Mengdi Wang · 2021
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Later among the works it cites.
Should i run offline reinforcement learning or behavioral cloning?
Aviral Kumar, Joey Hong, Anikait Singh, and Sergey Levine · 2021
Later among the works it cites.
Model selection in batch policy optimization
Jonathan N Lee, George Tucker, Ofir Nachum, and Bo Dai · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Later among the works it cites.
Masatoshi Uehara, Masaaki Imaizumi, Nan Jiang, Nathan Kallus, Wen Sun, and Tengyang Xie · 2021
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage, 2021
Masatoshi Uehara and Wen Sun · 2021
Later among the works it cites.
Minimax model learning
Cameron Voloshin, Nan Jiang, and Yisong Yue · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal · 2021
Later among the works it cites.
Pessimistic model selection for offline deep reinforcement learning, 2021
Chao-Han Huck Yang, Zhengling Qi, Yifan Cui, and Pin-Yu Chen · 2021
Later among the works it cites.
Towards instance-optimal offline reinforcement learning with pessimism
Ming Yin and Yu-Xiang Wang · 2021
Later among the works it cites.
Almost optimal batch-regret tradeoff for batch linear contextual bandits, 2021
Zihan Zhang, Xiangyang Ji, and Yuan Zhou · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Andrea Zanette, Martin J Wainwright, and Emma Brunskill · 2021
Later among the works it cites.
Xiaohong Chen and Zhengling Qi · 2022
Closest in time.
Offline reinforcement learning with realizability and single-policy concentrability
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, and Jason D Lee · 2022
Closest in time.
Efficient reinforcement learning in block MDPs: A model-free representation learning approach, 2022
Xuezhou Zhang, Yuda Song, Masatoshi Uehara, Mengdi Wang, Alekh Agarwal, and Wen Sun · 2022
Closest in time.