Fetching the paper…
Reading the bibliography…
In online learning problems, exploiting low variance plays an important role in obtaining tight performance guarantees yet is challenging because variances are often not known a priori.
Empirical processes: theory and applications
D. Pollard · 1990
Earlier work this paper cites.
Using Confidence Bounds for Exploitation-Exploration Trade-offs
P. Auer · 2002
Earlier work this paper cites.
The Nonstochastic Multiarmed Bandit Problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 2003
Earlier work this paper cites.
Provably Efficient Reinforcement Learning with General Value Function Approximation
R. Wang, R. Salakhutdinov, and L. F. Yang · 2005
Earlier work this paper cites.
Use of variance estimation in the multi-armed bandit problem
J.-Y. Audibert, R. Munos, and C. Szepesvari · 2006
Earlier work this paper cites.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Stochastic Linear Optimization under Bandit Feedback
V. Dani, T. P. Hayes, and S. M. Kakade · 2008
Earlier work this paper cites.
Extracting certainty from uncertainty: Regret bounded by variation in costs
E. Hazan and S. Kale · 2010
Earlier work this paper cites.
Improved Algorithms for Linear Stochastic Bandits
Y. Abbasi-Yadkori, D. Pal, and C. Szepesvari · 2011
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
D. Russo and B. Van Roy · 2013
Earlier work this paper cites.
Efficient exploration and value function generalization in deterministic systems
Z. Wen and B. Van Roy · 2013
Cited alongside, same era.
Pac reinforcement learning with rich observations
A. Krishnamurthy, A. Agarwal, and J. Langford · 2016
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
N. Jiang, A. Krishnamurthy, A. Agarwal, J. Langford, and R. E. Schapire · 2017
Cited alongside, same era.
On Oracle-Efficient PAC RL with Rich Observations
C. Dann, N. Jiang, A. Krishnamurthy, A. Agarwal, J. Langford, and R. E. Schapire · 2018
Cited alongside, same era.
Provably Efficient Q-learning with Function Approximation via Distribution Shift Error Checking Oracle
S. S. Du, Y. Luo, R. Wang, and H. Zhang · 2019
Cited alongside, same era.
Nearly Minimax-Optimal Regret for Linearly Parameterized Bandits
Y. Li, Y. Wang, and Y. Zhou · 2019
Provably Efficient Exploration for Reinforcement Learning Using Unsupervised Learning
F. Feng, R. Wang, W. Yin, S. S. Du, and L. Yang · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
C. Jin, Z. Yang, Z. Wang, and M. I. Jordan · 2020
Later among the works it cites.
Bandit Algorithms
T. Lattimore and C. Szepesvári · 2020
Later among the works it cites.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
D. Misra, M. Henaff, A. Krishnamurthy, and J. Langford · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
L. Yang and M. Wang · 2020
Later among the works it cites.
Learning near optimal policies with low inherent bellman error
A. Zanette, A. Lazaric, M. Kochenderfer, and E. Brunskill · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
W. Sun, N. Jiang, A. Krishnamurthy, A. Agarwal, and J. Langford · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
L. Yang and M. Wang · 2019
Cited alongside, same era.
Agnostic q q -learning with function approximation in deterministic systems: Near-optimal bounds on approximation error and sample complexity
S. S. Du, J. D. Lee, G. Mahajan, and R. Wang · 2020
Cited alongside, same era.
On Reward-Free Reinforcement Learning with Linear Function Approximation
R. Wang, S. S. Du, L. Yang, and R. R. Salakhutdinov
Cited in the paper.
Optimism in reinforcement learning with generalized linear function approximation
Y. Wang, R. Wang, S. S. Du, and A. Krishnamurthy
Cited in the paper.
Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon
Z. Zhang, X. Ji, and S. Du
Cited in the paper.
Later among the works it cites.
Logarithmic regret for reinforcement learning with linear function approximation
J. He, D. Zhou, and Q. Gu · 2021
Closest in time.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
D. Zhou, Q. Gu, and C. Szepesvari · 2021
Closest in time.
First-order regret in reinforcement learning with linear function approximation: A robust estimation approach
A. J. Wagenmaker, Y. Chen, M. Simchowitz, S. Du, and K. Jamieson · 2022
Closest in time.