Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) aims to find an optimal policy for Markov decision processes (MDPs) using a pre-collected dataset.
Error bounds and convergence analysis of feasible descent methods: a general approach
Zhi-Quan Luo and Paul Tseng · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman · 1994
Earlier work this paper cites.
Error bounds in mathematical programming
Jong-Shi Pang · 1997
Earlier work this paper cites.
Empirical Processes in M-estimation , volume 6
Sara A Geer · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
The linear programming approach to approximate dynamic programming
Daniela Pucci De Farias and Benjamin Van Roy · 2003
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Approximate policy iteration schemes: A comparison
Bruno Scherrer · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Dynamic programming and optimal control (4th edition)
Dimitri P Bertsekas · 2017
Earlier work this paper cites.
1 year, 1000 km: The oxford robotcar dataset
Will Maddern, Geoffrey Pascoe, Chris Linegar, and Paul Newman · 2017
Earlier work this paper cites.
Off-policy evaluation for slate recommendation
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miro Dudik, John Langford, Damien Jose, and Imed Zitouni · 2017
Earlier work this paper cites.
Deep reinforcement learning for automated radiation adaptation in lung cancer
Huan-Hsin Tseng, Yi Luo, Sunan Cui, Jen-Tzung Chien, Randall K Ten Haken, and Issam El Naqa · 2017
Earlier work this paper cites.
Starcraft II: A new challenge for reinforcement learning
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, Julian Schrittwieser, et al · 2017
Earlier work this paper cites.
Mengdi Wang · 2017
Earlier work this paper cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Sergey Levine, Peter Pastor, Alex Krizhevsky, Julian Ibarz, and Deirdre Quillen · 2018
Earlier work this paper cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Cited alongside, same era.
Near-optimal time and sample complexities for solving Markov decision processes with a generative model
Aaron Sidford, Mengdi Wang, Xian Wu, Lin Yang, and Yinyu Ye · 2018
Cited alongside, same era.
Politex: Regret bounds for policy iteration using expert prediction
Yasin Abbasi-Yadkori, Peter Bartlett, Kush Bhatia, Nevena Lazic, Csaba Szepesvari, and Gellért Weisz · 2019
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Offline reinforcement learning: Fundamental barriers for value function approximation
Dylan J Foster, Akshay Krishnamurthy, David Simchi-Levi, and Yunzong Xu · 2021
Later among the works it cites.
Towards tight bounds on the sample complexity of average-reward MDPs
Yujia Jin and Aaron Sidford · 2021
Later among the works it cites.
Optidice: Offline policy optimization via stationary distribution correction estimation
Jongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau, and Kee-Eung Kim · 2021
Later among the works it cites.
Reinforcement learning in linear MDPs: Constant regret and representation selection
Matteo Papini, Andrea Tirinzoni, Aldo Pacchiano, Marcello Restelli, Alessandro Lazaric, and Matteo Pirotta · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular MDPs
Max Simchowitz and Kevin G Jamieson · 2019
Cited alongside, same era.
Doubly robust bias reduction in infinite horizon off-policy estimation
Ziyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou, and Qiang Liu · 2019
Cited alongside, same era.
A variant of the Wang-Foster-Kakade lower bound for the discounted setting
Philip Amortila, Nan Jiang, and Tengyang Xie · 2020
Cited alongside, same era.
Minimax value interval for off-policy evaluation and policy optimization
Nan Jiang and Jiawei Huang · 2020
Cited alongside, same era.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2020
Cited alongside, same era.
Efficiently solving MDPs with stochastic mirror descent
Yujia Jin and Aaron Sidford · 2020
Cited alongside, same era.
Later among the works it cites.
Pessimistic model-based offline RL: Pac bounds and posterior sampling under partial coverage
Masatoshi Uehara and Wen Sun · 2021
Later among the works it cites.
Batch value-function approximation with only realizability
Tengyang Xie and Nan Jiang · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal · 2021
Later among the works it cites.
Q-learning with logarithmic regret
Kunhe Yang, Lin Yang, and Simon Du · 2021
Later among the works it cites.
Exponential lower bounds for batch reinforcement learning: Batch RL can be exponentially harder than online RL
Andrea Zanette · 2021
Later among the works it cites.
Finite-sample analysis for decentralized batch multiagent reinforcement learning with networked agents
Kaiqing Zhang, Zhuoran Yang, Han Liu, Tong Zhang, and Tamer Başar · 2021
Later among the works it cites.
Offline reinforcement learning under value and density-ratio realizability: the power of gaps
Jinglin Chen and Nan Jiang · 2022
Closest in time.
Adversarially trained actor critic for offline reinforcement learning
Ching-An Cheng, Tengyang Xie, Nan Jiang, and Alekh Agarwal · 2022
Closest in time.
What is a good metric to study generalization of minimax learners?
Asuman E. Ozdaglar, Sarath Pattathil, Jiawei Zhang, and Kaiqing Zhang · 2022
Closest in time.
Near sample-optimal reduction-based policy learning for average reward MDP
Jinghan Wang, Mengdi Wang, and Lin F Yang · 2022
Closest in time.
Offline reinforcement learning with realizability and single-policy concentrability
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, and Jason Lee · 2022
Closest in time.
Optimal conservative offline RL with general function approximation via augmented lagrangian
Paria Rashidinejad, Hanlin Zhu, Kunhe Yang, Stuart Russell, and Jiantao Jiao · 2023
Closest in time.
Offline primal-dual reinforcement learning for linear MDPs
Germano Gabbianelli, Gergely Neu, Matteo Papini, and Nneka M Okolo · 2024
Closest in time.
Importance weighted actor-critic for optimal conservative offline reinforcement learning
Hanlin Zhu, Paria Rashidinejad, and Jiantao Jiao · 2024
Closest in time.