Fetching the paper…
Reading the bibliography…
This paper studies regret minimization with randomized value functions in reinforcement learning.
Frequentist regret bounds for randomized least-squares value iteration
Zanette, A.; Brandfonbrener, D.; Brunskill, E.; Pirotta, M.; and Lazaric, A. 2020 · 1964
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G.; and Parr, R. 2003 · 2003
Earlier work this paper cites.
Extremely randomized trees
Geurts, P.; Ernst, D.; and Wehenkel, L. 2006 · 2006
Earlier work this paper cites.
On Optimism in Model-Based Reinforcement Learning
Pacchiano, A.; Ball, P.; Parker-Holder, J.; Choromanski, K.; and Roberts, S. 2020 · 2006
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T.; Ortner, R.; and Auer, P. 2010 · 2010
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Chapelle, O.; and Li, L. 2011 · 2011
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S.; and Goyal, N. 2013 · 2013
Earlier work this paper cites.
(More) efficient reinforcement learning via posterior sampling
Osband, I.; Russo, D.; and Van Roy, B. 2013 · 2013
Earlier work this paper cites.
Evaluation and analysis of the performance of the EXP3 algorithm in stochastic environments
Seldin, Y.; Szepesvári, C.; Auer, P.; and Abbasi-Yadkori, Y. 2013 · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. 2014 · 2014
Earlier work this paper cites.
Deep exploration via bootstrapped DQN
Osband, I.; Blundell, C.; Pritzel, A.; and Van Roy, B. 2016 · 2016
Earlier work this paper cites.
Generalization and exploration via randomized value functions
Osband, I.; Van Roy, B.; and Wen, Z. 2016 · 2016
Earlier work this paper cites.
Linear thompson sampling revisited
Abeille, M.; Lazaric, A.; et al. 2017 · 2017
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Agrawal, S.; and Jia, R. 2017 · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G.; Osband, I.; and Munos, R. 2017 · 2017
Cited alongside, same era.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Dann, C.; Lattimore, T.; and Brunskill, E. 2017 · 2017
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N.; Krishnamurthy, A.; Agarwal, A.; Langford, J.; and Schapire, R. E. 2017 · 2017
Cited alongside, same era.
Why is posterior sampling better than optimism for reinforcement learning?
Osband, I.; and Van Roy, B. 2017 · 2017
Explicit explore-exploit algorithms in continuous state spaces
Henaff, M. 2019 · 2019
Later among the works it cites.
Garbage in, reward out: Bootstrapping exploration in multi-armed bandits
Kveton, B.; Szepesvari, C.; Vaswani, S.; Wen, Z.; Lattimore, T.; and Ghavamzadeh, M. 2019 · 2019
Later among the works it cites.
Deep Exploration via Randomized Value Functions
Osband, I.; Van Roy, B.; Russo, D. J.; and Wen, Z. 2019 · 2019
Later among the works it cites.
Worst-case regret bounds for exploration via randomized value functions
Russo, D. 2019 · 2019
Later among the works it cites.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Sun, W.; Jiang, N.; Krishnamurthy, A.; Agarwal, A.; and Langford, J. 2019 · 2019
Later among the works it cites.
Basic tail and concentration bounds , 21–57
Wainwright, M. J. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Noisy Networks For Exploration
Fortunato, M.; Azar, M. G.; Piot, B.; Menick, J.; Hessel, M.; Osband, I.; Graves, A.; Mnih, V.; Munos, R.; Hassabis, D.; et al. 2018 · 2018
Cited alongside, same era.
Is q-learning provably efficient?
Jin, C.; Allen-Zhu, Z.; Bubeck, S.; and Jordan, M. I. 2018 · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
Osband, I.; Aslanides, J.; and Cassirer, A. 2018 · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Cited alongside, same era.
Information-Theoretic Considerations in Batch Reinforcement Learning
Chen, J.; and Jiang, N. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Worst-Case Regret Bound for Perturbation Based Exploration in Reinforcement Learning
Xu, Z.; and Tewari, A. 2019 · 2019
Later among the works it cites.
Tighter Problem-Dependent Regret Bounds in Reinforcement Learning without Domain Knowledge using Value Function Bounds
Zanette, A.; and Brunskill, E. 2019 · 2019
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C.; Yang, Z.; Wang, Z.; and Jordan, M. I. 2020 · 2020
Closest in time.
Bandit algorithms
Lattimore, T.; and Szepesvári, C. 2020 · 2020
Closest in time.
Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret Bound
Yang, L. F.; and Wang, M. 2020 · 2020
Closest in time.