Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL), where the agent aims to learn the optimal policy based on the data collected by a behavior policy, has attracted increasing attention in recent years.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 1964
Earlier work this paper cites.
Prediction, learning, and games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
What are the statistical limits of offline rl with linear function approximation?
Ruosong Wang, Dean P Foster, and Sham M Kakade · 2010
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
Active learning for cost-sensitive classification
Akshay Krishnamurthy, Alekh Agarwal, Tzu-Kuo Huang, Hal Daumé III, and John Langford · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
On oracle-efficient pac rl with rich observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2018
Earlier work this paper cites.
Practical contextual bandits with regression oracles
Dylan Foster, Alekh Agarwal, Miroslav Dudík, Haipeng Luo, and Robert Schapire · 2018
Earlier work this paper cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Sergey Levine, Peter Pastor, Alex Krizhevsky, Julian Ibarz, and Deirdre Quillen · 2018
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Earlier work this paper cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S Du, Sham M Kakade, Ruosong Wang, and Lin F Yang · 2019
Earlier work this paper cites.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Earlier work this paper cites.
Sample-optimal parametric q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Earlier work this paper cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Earlier work this paper cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Aditya Modi, Nan Jiang, Ambuj Tewari, and Satinder Singh · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Cited alongside, same era.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin Yang and Mengdi Wang · 2020
Cited alongside, same era.
Bilinear classes: A structural framework for provable generalization in rl
Simon Du, Sham Kakade, Jason Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Cited alongside, same era.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Cited alongside, same era.
Magnetic control of tokamak plasmas through deep reinforcement learning
Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de Las Casas, et al · 2022
Later among the works it cites.
Achieving minimax rates in pool-based batch active learning
Claudio Gentile, Zhilei Wang, and Tong Zhang · 2022
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear markov decision processes
Jiafan He, Heyang Zhao, Dongruo Zhou, and Quanquan Gu · 2022
Later among the works it cites.
Nearly minimax optimal reinforcement learning with linear function approximation
Pihe Hu, Yu Chen, and Longbo Huang · 2022
Later among the works it cites.
Settling the sample complexity of model-based offline reinforcement learning
Gen Li, Laixi Shi, Yuxin Chen, Yuejie Chi, and Yuting Wei · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Logarithmic regret for reinforcement learning with linear function approximation
Jiafan He, Dongruo Zhou, and Quanquan Gu · 2021
Cited alongside, same era.
Online sub-sampling for reinforcement learning with general function approximation
Dingwen Kong, Ruslan Salakhutdinov, Ruosong Wang, and Lin F Yang · 2021
Cited alongside, same era.
Variance-aware off-policy evaluation with linear function approximation
Yifei Min, Tianhao Wang, Dongruo Zhou, and Quanquan Gu · 2021
Cited alongside, same era.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Cited alongside, same era.
Pessimistic model-based offline reinforcement learning under partial coverage
Masatoshi Uehara and Wen Sun · 2021
Cited alongside, same era.
Exponential lower bounds for planning in mdps with linearly-realizable optimal action-value functions
Gellért Weisz, Philip Amortila, and Csaba Szepesvári · 2021
Cited alongside, same era.
Towards instance-optimal offline reinforcement learning with pessimism
Ming Yin and Yu-Xiang Wang · 2021
Cited alongside, same era.
Later among the works it cites.
Optimal conservative offline rl with general function approximation via augmented lagrangian
Paria Rashidinejad, Hanlin Zhu, Kunhe Yang, Stuart Russell, and Jiantao Jiao · 2022
Later among the works it cites.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Laixi Shi, Gen Li, Yuting Wei, Yuxin Chen, and Yuejie Chi · 2022
Later among the works it cites.
Offline reinforcement learning with realizability and single-policy concentrability
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, and Jason Lee · 2022
Later among the works it cites.
Pessimistic minimax value iteration: Provably efficient equilibrium learning from offline datasets
Han Zhong, Wei Xiong, Jiyuan Tan, Liwei Wang, Tong Zhang, Zhaoran Wang, and Zhuoran Yang · 2022
Later among the works it cites.
Vo q q l: Towards optimal regret in model-free rl with nonlinear function approximation
Alekh Agarwal, Yujia Jin, and Tong Zhang · 2023
Closest in time.
Viper: Provably efficient algorithm for offline rl with neural function approximation
Thanh Nguyen-Tang and Raman Arora · 2023
Closest in time.
On instance-dependent bounds for offline reinforcement learning with linear function approximation
Thanh Nguyen-Tang, Ming Yin, Sunil Gupta, Svetha Venkatesh, and Raman Arora · 2023
Closest in time.
Revisiting the linear-programming framework for offline rl with general function approximation
Asuman E Ozdaglar, Sarath Pattathil, Jiawei Zhang, and Kaiqing Zhang · 2023
Closest in time.
Nearly minimax optimal offline reinforcement learning with linear function approximation: Single-agent mdp and markov game
Wei Xiong, Han Zhong, Chengshuai Shi, Cong Shen, Liwei Wang, and Tong Zhang · 2023
Closest in time.
Corruption-robust algorithms with uncertainty weighting for nonlinear contextual bandits and markov decision processes
Chenlu Ye, Wei Xiong, Quanquan Gu, and Tong Zhang · 2023
Closest in time.
Importance weighted actor-critic for optimal conservative offline reinforcement learning
Hanlin Zhu, Paria Rashidinejad, and Jiantao Jiao · 2023
Closest in time.