Fetching the paper…
Reading the bibliography…
Recent theoretical work studies sample-efficient reinforcement learning (RL) extensively in two settings: learning interactively in the environment (online RL), or learning from an offline dataset (offline RL).
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
Error bounds for approximate policy iteration
R. Munos · 2003
Earlier work this paper cites.
Finite time bounds for sampling based fitted value iteration
C. Szepesvári and R. Munos · 2005
Earlier work this paper cites.
Theory of point estimation
E. L. Lehmann and G. Casella · 2006
Earlier work this paper cites.
Provably good batch reinforcement learning without great exploration
Y. Liu, A. Swaminathan, A. Agarwal, and E. Brunskill · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
A. Antos, C. Szepesvári, and R. Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
R. Munos and C. Szepesvári · 2008
Earlier work this paper cites.
Provably efficient reward-agnostic navigation with linear value iteration
A. Zanette, A. Lazaric, M. J. Kochenderfer, and E. Brunskill · 2008
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
A. Maurer and M. Pontil · 2009
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
A. M. Farahmand, R. Munos, and C. Szepesvári · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
A sharp analysis of model-based reinforcement learning with self-play
Q. Liu, T. Yu, Y. Bai, and C. Jin · 2010
Earlier work this paper cites.
Is pessimism provably efficient for offline rl?
Y. Jin, Z. Yang, and Z. Wang · 2012
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
I. Osband and B. V. Roy · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
S. Agrawal and R. Jia · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
M. G. Azar, I. Osband, and R. Munos · 2017
Earlier work this paper cites.
Unifying pac and regret: uniform pac bounds for episodic reinforcement learning
C. Dann, T. Lattimore, and E. Brunskill · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
N. Jiang, A. Krishnamurthy, A. Agarwal, J. Langford, and R. E. Schapire · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Cited alongside, same era.
Boosted fitted q-iteration
S. Tosatto, M. Pirotta, C. d’Eramo, and M. Restelli · 2017
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Cited alongside, same era.
Is q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Cited alongside, same era.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Bandit algorithms
T. Lattimore and C. Szepesvári · 2020
Later among the works it cites.
Learning quadrupedal locomotion over challenging terrain
J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter · 2020
Later among the works it cites.
Deployment-efficient reinforcement learning via model-based offline optimization
T. Matsushima, H. Furuta, Y. Matsuo, O. Nachum, and S. Gu · 2020
Later among the works it cites.
On bonus based exploration methods in the arcade learning environment
A. A. Taiga, W. Fedus, M. C. Machado, A. Courville, and M. G. Bellemare · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, et al · 2019
Cited alongside, same era.
Provably efficient q-learning with low switching cost
Y. Bai, T. Xie, N. Jiang, and Y.-X. Wang · 2019
Cited alongside, same era.
Emergent tool use from multi-agent autocurricula
B. Baker, I. Kanitscheider, T. Markov, Y. Wu, G. Powell, B. McGrew, and I. Mordatch · 2019
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
J. Chen and N. Jiang · 2019
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
C. Dann, L. Li, W. Wei, and E. Brunskill · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Cited alongside, same era.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
R. Wang, R. R. Salakhutdinov, and L. Yang · 2020
Later among the works it cites.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
T. Xie and N. Jiang · 2020
Later among the works it cites.
Z. Yang, C. Jin, Z. Wang, M. Wang, and M. I. Jordan · 2020
Later among the works it cites.
Near optimal provable uniform convergence in off-policy evaluation for reinforcement learning
M. Yin, Y. Bai, and Y.-X. Wang · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
Almost optimal model-free reinforcement learningvia reference-advantage decomposition
Z. Zhang, Y. Zhou, and X. Ji · 2020
Later among the works it cites.
Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited
O. D. Domingues, P. Ménard, E. Kaufmann, and M. Valko · 2021
Closest in time.
Bilinear classes: A structural framework for provable generalization in rl
S. S. Du, S. M. Kakade, J. D. Lee, S. Lovett, G. Mahajan, W. Sun, and R. Wang · 2021
Closest in time.
A provably efficient algorithm for linear markov decision process with low switching cost
M. Gao, T. Xie, S. S. Du, and L. F. Yang · 2021
Closest in time.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
C. Jin, Q. Liu, and S. Miryoosefi · 2021
Closest in time.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
P. Rashidinejad, B. Zhu, C. Ma, J. Jiao, and S. Russell · 2021
Closest in time.
Nearly horizon-free offline reinforcement learning
T. Ren, J. Li, B. Dai, S. S. Du, and S. Sanghavi · 2021
Closest in time.
D. Su, J. D. Lee, J. M. Mulvey, and H. V. Poor · 2021
Closest in time.
T. Wang, D. Zhou, and Q. Gu · 2021
Closest in time.
Near-optimal offline reinforcement learning via double variance reduction
M. Yin, Y. Bai, and Y.-X. Wang · 2021
Closest in time.