Fetching the paper…
Reading the bibliography…
Two central paradigms have emerged in the reinforcement learning (RL) community: online RL and offline RL.
Frequentist regret bounds for randomized least-squares value iteration
Zanette, A., Brandfonbrener, D., Brunskill, E., Pirotta, M., and Lazaric, A · 1964
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M · 2003
Earlier work this paper cites.
Optimistic linear programming gives logarithmic regret for irreducible mdps
Tewari, A. and Bartlett, P · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Auer, P., Jaksch, T., and Ortner, R · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
Introduction to nonparametric estimation., 2009
Tsybakov, A. B · 2009
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Zhou, D., He, J., and Gu, Q · 2011
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Ross, S. and Bagnell, J. A · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J · 2013
Earlier work this paper cites.
On the complexity of bandit and derivative-free stochastic convex optimization
Shamir, O · 2013
Earlier work this paper cites.
On the complexity of best-arm identification in multi-armed bandit models
Kaufmann, E., Cappé, O., and Garivier, A · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Agrawal, S. and Jia, R · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Earlier work this paper cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E · 2017
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2017
Earlier work this paper cites.
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., et al · 2018
Earlier work this paper cites.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Earlier work this paper cites.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Earlier work this paper cites.
Exploration in structured reinforcement learning
Ok, J., Proutiere, A., and Tranos, D · 2018
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J. and Jiang, N · 2019
Earlier work this paper cites.
Sequential experimental design for transductive linear bandits
Fiez, T., Jain, L., Jamieson, K. G., and Ratliff, L · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M. and Jamieson, K. G · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Towards tractable optimism in model-based reinforcement learning
Pacchiano, A., Ball, P., Parker-Holder, J., Choromanski, K., and Roberts, S · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S · 2021
Later among the works it cites.
Bandits with partially observable confounded data
Tennenholtz, G., Shalit, U., Mannor, S., and Efroni, Y · 2021
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M. and Sun, W · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zanette, A. and Brunskill, E · 2019
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Ayoub, A., Jia, Z., Szepesvari, C., Wang, M., and Yang, L · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Logarithmic regret for reinforcement learning with linear function approximation
He, J., Zhou, D., and Gu, Q · 2020
Cited alongside, same era.
Minimax value interval for off-policy evaluation and policy optimization
Jiang, N. and Huang, J · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2020
Cited alongside, same era.
The true sample complexity of identifying good arms
Katz-Samuels, J. and Jamieson, K · 2020
Cited alongside, same era.
Weisz, G., Amortila, P., and Szepesvári, C · 2021
Later among the works it cites.
Batch value-function approximation with only realizability
Xie, T. and Jiang, N · 2021
Later among the works it cites.
Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap
Xu, H., Ma, T., and Du, S · 2021
Later among the works it cites.
Q-learning with logarithmic regret
Yang, K., Yang, L., and Du, S · 2021
Later among the works it cites.
Near-optimal offline reinforcement learning via double variance reduction
Yin, M., Bai, Y., and Wang, Y.-X · 2021
Later among the works it cites.
Reinforcement learning in healthcare: A survey
Yu, C., Liu, J., Nemati, S., and Yin, G · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Zanette, A., Wainwright, M. J., and Brunskill, E · 2021
Later among the works it cites.
Offline reinforcement learning under value and density-ratio realizability: the power of gaps
Chen, J. and Jiang, N · 2022
Closest in time.
A free lunch from the noise: Provable and practical exploration for representation learning
Ren, T., Zhang, T., Szepesvári, C., and Dai, B · 2022
Closest in time.
Hybrid rl: Using both offline and online data can make rl efficient
Song, Y., Zhou, Y., Sekhari, A., Bagnell, J. A., Krishnamurthy, A., and Sun, W · 2022
Closest in time.
Optimistic pac reinforcement learning: the instance-dependent view
Tirinzoni, A., Al-Marjani, A., and Kaufmann, E · 2022
Closest in time.
Instance-dependent near-optimal policy identification in linear mdps via online experiment design
Wagenmaker, A. and Jamieson, K · 2022
Closest in time.
The role of coverage in online reinforcement learning
Xie, T., Foster, D. J., Bai, Y., Jiang, N., and Kakade, S. M · 2022
Closest in time.
Yin, M., Duan, Y., Wang, M., and Wang, Y.-X · 2022
Closest in time.
Offline reinforcement learning with realizability and single-policy concentrability
Zhan, W., Huang, B., Huang, A., Jiang, N., and Lee, J · 2022
Closest in time.
Efficient online reinforcement learning with offline data
Ball, P. J., Smith, L., Kostrikov, I., and Levine, S · 2023
Closest in time.
Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning
Nakamoto, M., Zhai, Y., Singh, A., Mark, M. S., Ma, Y., Finn, C., Kumar, A., and Levine, S · 2023
Closest in time.
Instance-optimality in interactive decision making: Toward a non-asymptotic theory
Wagenmaker, A. and Foster, D. J · 2023
Closest in time.
Adaptive policy learning for offline-to-online reinforcement learning
Zheng, H., Luo, X., Wei, P., Song, X., Li, D., and Jiang, J · 2023
Closest in time.