Fetching the paper…
Reading the bibliography…
Recently, several studies (Zhou et al., 2021a; Zhang et al., 2021b; Kim et al., 2021; Zhou and Gu, 2022) have provided variance-dependent regret bounds for linear contextual bandits, which interpolates the regret for the worst-case regime and the deterministic reward regime.
Is a good representation sufficient for sample efficient reinforcement learning?
Du, S. S · 1910
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Robbins, H · 1952
Earlier work this paper cites.
On tail probabilities for martingales
Freedman, D. A · 1975
Earlier work this paper cites.
Finite-time regret bounds for the multiarmed bandit problem
Cesa-Bianchi, N · 1998
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P · 2002
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
Abe, N · 2003
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N · 2006
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V · 2008
Earlier work this paper cites.
The elliptical potential lemma revisited
Carpentier, A · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Chu, W · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S · 2012
Earlier work this paper cites.
Matrix computations
Golub, G. H · 2013
Earlier work this paper cites.
How hard is my mdp?” the distribution-norm to the rescue”
Maillard, O.-A · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Cited alongside, same era.
Linear multi-resource allocation with semi-bandit feedback
Lattimore, T · 2015
Cited alongside, same era.
Multi-armed bandit models for the optimal design of clinical trials: benefits and challenges
Villar, S. S · 2015
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N · 2017
Cited alongside, same era.
On oracle-efficient pac rl with rich observations
Dann, C · 2018
Cited alongside, same era.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Yang, L · 2020
Later among the works it cites.
Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning
Dann, C · 2021
Later among the works it cites.
Improved regret analysis for variance-adaptive linear bandits and horizon-free linear mixture mdps
Kim, Y · 2021
Later among the works it cites.
Settling the horizon-dependence of sample complexity in reinforcement learning
Li, Y · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Open problem: The dependence of sample complexity lower bounds on planning horizon
Jiang, N · 2018
Cited alongside, same era.
Information directed sampling and bandits with heteroscedastic noise
Kirschner, J · 2018
Cited alongside, same era.
Nearly minimax-optimal regret for linearly parameterized bandits
Li, Y · 2019
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M · 2019
Cited alongside, same era.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Sun, W · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Yang, L · 2019
Cited alongside, same era.
Weisz, G · 2021
Later among the works it cites.
Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap
Xu, H · 2021
Later among the works it cites.
Vo q q l: Towards optimal regret in model-free rl with nonlinear function approximation
Agarwal, A · 2022
Later among the works it cites.
Variance-aware sparse linear bandits
Dai, Y · 2022
Later among the works it cites.
Nearly minimax optimal reinforcement learning with linear function approximation
Hu, P · 2022
Later among the works it cites.
First-order regret in reinforcement learning with linear function approximation: A robust estimation approach
Wagenmaker, A. J · 2022
Later among the works it cites.
Horizon-free reinforcement learning in polynomial time: the power of stationary policies
Zhang, Z · 2022
Later among the works it cites.
Zhao, H · 2022
Later among the works it cites.
Computationally efficient horizon-free reinforcement learning for linear mixture mdps
Zhou, D · 2022
Later among the works it cites.