Fetching the paper…
Reading the bibliography…
In this paper, we study the problem of efficient online reinforcement learning in the infinite horizon setting when there is an offline dataset to start with.
Dynamic programming and optimal control: Volume II , volume 2
D. Bertsekas · 1995
Earlier work this paper cites.
Learning from demonstration
S. Schaal · 1996
Earlier work this paper cites.
Inequalities for the L1 deviation of the empirical distribution
T. Weissman, E. Ordentlich, G. Seroussi, S. Verdu, and M. J. Weinberger · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
D. Ernst, P. Geurts, and L. Wehenkel · 2005
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
P. Auer, T. Jaksch, and R. Ortner · 2008
Earlier work this paper cites.
(More) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. Van Roy · 2013
Earlier work this paper cites.
An information-theoretic analysis of Thompson sampling
D. Russo and B. Van Roy · 2016
Earlier work this paper cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
S. Agrawal and R. Jia · 2017
Earlier work this paper cites.
Markov chains and mixing times , volume 107
D. A. Levin and Y. Peres · 2017
Earlier work this paper cites.
Learning unknown Markov decision processes: A Thompson sampling approach
Y. Ouyang, M. Gagrani, A. Nayyar, and R. Jain · 2017
Earlier work this paper cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. Riedmiller · 2017
Earlier work this paper cites.
Deep Q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, et al · 2018
Earlier work this paper cites.
A tutorial on Thompson sampling
D. J. Russo, B. Van Roy, A. Kazerouni, I. Osband, Z. Wen, et al · 2018
Earlier work this paper cites.
Deep exploration via randomized value functions
I. Osband, B. Van Roy, D. J. Russo, Z. Wen, et al · 2019
Earlier work this paper cites.
A. Argenson and G. Dulac-Arnold · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Conservative Q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Cited alongside, same era.
AWAC: Accelerating online reinforcement learning with offline datasets
A. Nair, A. Gupta, M. Dalal, and S. Levine · 2020
Cited alongside, same era.
MoDem: Accelerating visual model-based reinforcement learning with demonstrations
N. Hansen, Y. Lin, H. Su, X. Wang, V. Kumar, and A. Rajeswaran · 2022
Later among the works it cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Later among the works it cites.
When should we prefer offline reinforcement learning over behavioral cloning?
A. Kumar, J. Hong, A. Singh, and S. Levine · 2022
Later among the works it cites.
Offline-to-online reinforcement learning via balanced replay and pessimistic Q-ensemble
S. Lee, Y. Seo, K. Lee, P. Abbeel, and J. Shin · 2022
Later among the works it cites.
Hybrid RL: Using both offline and online data can make RL efficient
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Is pessimism provably efficient for offline RL?
Y. Jin, Z. Yang, and Z. Wang · 2021
Cited alongside, same era.
Offline reinforcement learning with implicit Q-learning
I. Kostrikov, A. Nair, and S. Levine · 2021
Cited alongside, same era.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
P. Rashidinejad, B. Zhu, C. Ma, J. Jiao, and S. Russell · 2021
Cited alongside, same era.
Online and offline reinforcement learning by planning with a learned model
J. Schrittwieser, T. Hubert, A. Mandhane, M. Barekatain, I. Antonoglou, and D. Silver · 2021
Cited alongside, same era.
Pessimistic model-based offline reinforcement learning under partial coverage
M. Uehara and W. Sun · 2021
Cited alongside, same era.
Model-based RL with optimistic posterior sampling: Structural conditions and sample complexity
A. Agarwal and T. Zhang · 2022
Cited alongside, same era.
Imitation learning by estimating expertise of demonstrators
M. Beliaev, A. Shih, S. Ermon, D. Sadigh, and R. Pedarsani · 2022
Cited alongside, same era.
Y. Song, Y. Zhou, A. Sekhari, J. A. Bagnell, A. Krishnamurthy, and W. Sun · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, et al · 2022
Later among the works it cites.
Leveraging offline data in online reinforcement learning
A. Wagenmaker and A. Pacchiano · 2022
Later among the works it cites.
Safe exploration for efficient policy evaluation and comparison
R. Wan, B. Kveton, and R. Song · 2022
Later among the works it cites.
Efficient online reinforcement learning with offline data
P. J. Ball, L. Smith, I. Kostrikov, and S. Levine · 2023
Closest in time.
Finetuning offline world models in the real world
Y. Feng, N. Hansen, Z. Xiong, C. Rajagopalan, and X. Wang · 2023
Closest in time.
Leveraging demonstrations to improve online learning: Quality matters
B. Hao, R. Jain, T. Lattimore, B. Van Roy, and Z. Wen · 2023
Closest in time.
Imitation bootstrapped reinforcement learning
H. Hu, S. Mirchandani, and D. Sadigh · 2023
Closest in time.
Adaptive policy learning for offline-to-online reinforcement learning
H. Zheng, X. Luo, P. Wei, X. Song, D. Li, and J. Jiang · 2023
Closest in time.