Fetching the paper…
Reading the bibliography…
This paper studies model-based reinforcement learning (RL) for regret minimization.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. F. and Wang, M · 1905
Earlier work this paper cites.
Introduction to queueing theory
Kovalenko, B. G. I. N · 1968
Earlier work this paper cites.
Stochastic optimal control: the discrete-time case
Bertsekas, D. P. and Shreve, S · 1978
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Neuro-dynamic programming
Bertsekas, D. P. and Tsitsiklis, J. N · 1996
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. N. and Van Roy, B · 1997
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Strens, M. J. A · 2000
Earlier work this paper cites.
Artificial Intelligence – a modern approach
Russel, S. and Norvig, P · 2003
Earlier work this paper cites.
Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?
Wang, R., Du, S. S., Yang, L. F., and Kakade, S. M · 2005
Earlier work this paper cites.
Provably efficient reinforcement learning with general value function approximation
Wang, R., Salakhutdinov, R., and Yang, L. F · 2005
Earlier work this paper cites.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
Parr, R., Li, L., Taylor, G., Painter-Wakefield, C., and Littman, M. L · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for Markov decision processes
Strehl, A. and Littman, M · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Rusmevichientong, P. and Tsitsiklis, J. N · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J · 2013
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Efficient exploration and value function generalization in deterministic systems
Wen, Z. and Van Roy, B · 2013
Cited alongside, same era.
Model-based reinforcement learning and the Eluder dimension
Osband, I. and Van Roy, B · 2014
Cited alongside, same era.
Generalization and exploration via randomized value functions
Osband, I., Van Roy, B., and Wen, Z · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
Russo, D. and Van Roy, B · 2014
Cited alongside, same era.
Bayesian optimal control of smoothly parameterized systems
Abbasi-Yadkori, Y. and Szepesvári, C · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Efficient reinforcement learning in deterministic systems with value function generalization
Wen, Z. and Van Roy, B · 2017
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
Dann, C., Li, L., Wei, W., and Brunskill, E · 2018
Later among the works it cites.
Is Q Q -learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Later among the works it cites.
Variance reduction methods for sublinear reinforcement learning
Kakade, S., Wang, M., and Yang, L. F · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
AlphaFold at CASP13
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Osband, I. and Van Roy, B · 2016
Cited alongside, same era.
Policy error bounds for model-based reinforcement learning with factored linear models
Pires, B. and Szepesvári, C · 2016
Cited alongside, same era.
Model-based reinforcement learning with parametrized physical models and optimism-driven exploration
Xie, C., Patil, S., Moldovan, T., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Agrawal, S. and Jia, R · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Cited alongside, same era.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E · 2017
Cited alongside, same era.
AlQuraishi, M · 2019
Later among the works it cites.
Alphastar: An evolutionary computation perspective
Arulkumaran, K., Cully, A., and Togelius, J · 2019
Later among the works it cites.
Du, S. S., Luo, Y., Wang, R., and Zhang, H · 2019
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2019
Later among the works it cites.
Model-based reinforcement learning for Atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., Mohiuddin, A., Sepassi, R., Tucker, G., and Michalewski, H · 2019
Later among the works it cites.
Learning with good feature representations in bandits and in RL with a generative model, 2019
Lattimore, T. and Szepesvári, C · 2019
Later among the works it cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A., Jiang, N., Tewari, A., and Singh, S · 2019
Later among the works it cites.
Worst-case regret bounds for exploration via randomized value functions
Russo, D · 2019
Later among the works it cites.
Mastering Atari, Go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., , Lillicrap, T., and Silver, D · 2019
Later among the works it cites.
Limiting extrapolation in linear approximate value iteration
Zanette, A., Lazaric, A., Kochenderfer, M. J., and Brunskill, E · 2019
Later among the works it cites.
Bandit Algorithms
Lattimore, T. and Szepesvári, C · 2020
Closest in time.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zhang, Z., Zhou, Y., and Ji, X · 2020
Closest in time.