Fetching the paper…
Reading the bibliography…
This paper presents new \emph{variance-aware} confidence sets for linear bandits and linear mixture Markov Decision Processes (MDPs).
Tight regret bounds for infinite-armed linear contextual bandits
Yingkai Li, Yining Wang, and Yuan Zhou · 1905
Earlier work this paper cites.
Weighted sums of certain dependent random variables
Kazuoki Azuma · 1967
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2002
Earlier work this paper cites.
Use of variance estimation in the multi-armed bandit problem
Jean-Yves Audibert, Rémi Munos, and Csaba Szepesvari · 2006
Earlier work this paper cites.
Model-free reinforcement learning: from clipped pseudo-regret to sample complexity
Zihan Zhang, Yuan Zhou, and Xiangyang Ji · 2006
Earlier work this paper cites.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Dongruo Zhou, Jiafan He, and Quanquan Gu · 2006
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
Empirical Bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Zihan Zhang, Xiangyang Ji, and Simon S Du · 2009
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L Bartlett and Ambuj Tewari · 2012
Earlier work this paper cites.
Pac bounds for discounted mdps
Tor Lattimore and Marcus Hutter · 2012
Earlier work this paper cites.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2012
Earlier work this paper cites.
Empirical bernstein inequality for martingales: Application to online learning
Thomas Peel, Sandrine Anthoine, and Liva Ralaivola · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Dan Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Efficient exploration and value function generalization in deterministic systems
Zheng Wen and Benjamin Van Roy · 2013
Earlier work this paper cites.
PAC reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
On oracle-efficient PAC-RL with rich observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E. Schapire · 2018
Cited alongside, same era.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Nan Jiang and Alekh Agarwal · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Improved optimistic algorithms for logistic bandits
Louis Faury, Marc Abeille, Clément Calauzènes, and Olivier Fercoq · 2020
Later among the works it cites.
Provably efficient exploration for RL with unsupervised learning
Fei Feng, Ruosong Wang, Wotao Yin, Simon S Du, and Lin F Yang · 2020
Later among the works it cites.
Logarithmic regret for reinforcement learning with linear function approximation
Jiafan He, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
Optimal best-arm identification in linear bandits
Yassir Jedra and Alexandre Proutiere · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2019
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2019
Cited alongside, same era.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2019
Cited alongside, same era.
Model-based RL in contextual decision processes: PAC bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Cited alongside, same era.
Optimism in reinforcement learning with generalized linear function approximation
Yining Wang, Ruosong Wang, Simon S Du, and Akshay Krishnamurthy · 2019
Cited alongside, same era.
Sample-optimal parametric Q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Cited alongside, same era.
An empirical process approach to the union bound: Practical algorithms for combinatorial and linear bandits
Julian Katz-Samuels, Lalit Jain, Kevin G Jamieson, et al · 2020
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Later among the works it cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Aditya Modi, Nan Jiang, Ambuj Tewari, and Satinder Singh · 2020
Later among the works it cites.
Efficient planning in large mdps with weak linear function approximation
Roshan Shariff and Csaba Szepesvári · 2020
Later among the works it cites.
Gellert Weisz, Philip Amortila, and Csaba Szepesvári · 2020
Later among the works it cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Lin F Yang and Mengdi Wang · 2020
Later among the works it cites.
Learning near optimal policies with low inherent bellman error
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
Instance-wise minimax-optimal algorithms for logistic bandits
Marc Abeille, Louis Faury, and Clément Calauzènes · 2021
Closest in time.
Bilinear classes: A structural framework for provable generalization in RL
Simon S Du, Sham M Kakade, Jason D Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Closest in time.
Bellman Eluder dimension: New rich classes of RL problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Closest in time.
Ucb momentum q-learning: Correcting the bias without forgetting
Pierre Menard, Omar Darwiche Domingues, Xuedong Shang, and Michal Valko · 2021
Closest in time.