Fetching the paper…
Reading the bibliography…
We consider model selection in stochastic bandit and reinforcement learning problems.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicoló Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Adaptive bandits: Towards the best history-dependent strategy
Odalric-Ambrym Maillard and Rémi Munos · 2011
Earlier work this paper cites.
Selecting the state-representation in reinforcement learning
Odalric-Ambrym Maillard, Daniil Ryabko, and Rémi Munos · 2011
Earlier work this paper cites.
Algorithms for Reinforcement Learning
Csaba Szepesvári · 2011
Earlier work this paper cites.
Optimal regret bounds for selecting the state representation in reinforcement learning
Odalric-Ambrym Maillard, Phuong Nguyen, Ronald Ortner, and Daniil Ryabko · 2013
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Benjamin Van Roy, and Daniel Russo · 2013
Cited alongside, same era.
Optimal regret bounds for selecting the state representation in reinforcement learning
Ronald Ortner, Odalric-Ambrym Maillard, and Daniil Ryabko · 2014
Cited alongside, same era.
Corralling a band of bandit algorithms
Alekh Agarwal, Haipeng Luo, Behnam Neyshabur, and Robert E. Schapire · 2017
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
POLITEX: Regret bounds for policy iteration using expert prediction
Yasin Abbasi-Yadkori, Peter L. Bartlett, Kush Bhatia, Nevena Lazic, Csaba Szepesvári, and Gellért Weisz · 2019
Model selection for contextual bandits
Dylan Foster, Akshay Krishnamurthy, and Haipeng Luo · 2019
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I. Jordan · 2019
Later among the works it cites.
Regret bounds for learning state representations in reinforcement learning
Ronald Ortner, Matteo Pirotta, Alessandro Lazaric, Ronan Fruit, and Odalric-Ambrym Maillard · 2019
Later among the works it cites.
Osom: A simultaneously optimal algorithm for multi-armed and linear contextual bandits
Niladri Chatterji, Vidya Muthukumar, and Peter L. Bartlett · 2020
Closest in time.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Aldo Pacchiano, My Phan, Yasin Abbasi-Yadkori, Anup Rao, Julian Zimmert, Tor Lattimore, and Csaba Szepesvari · 2020
Closest in time.