Fetching the paper…
Reading the bibliography…
Recent studies have shown that episodic reinforcement learning (RL) is not more difficult than contextual bandits, even with a long planning horizon and unknown state transitions.
Tight regret bounds for infinite-armed linear contextual bandits
Li, Y · 1905
Earlier work this paper cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. F · 1905
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Cai, Q · 1912
Earlier work this paper cites.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y · 1912
Earlier work this paper cites.
Weighted sums of certain dependent random variables
Azuma, K · 1967
Earlier work this paper cites.
On bernstein-type inequalities for martingales
Dzhaparidze, K · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Earlier work this paper cites.
Learning near optimal policies with low inherent Bellman error
Zanette, A · 2003
Earlier work this paper cites.
Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?
Wang, R · 2005
Earlier work this paper cites.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2006
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V · 2008
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L · 2010
Earlier work this paper cites.
Weisz, G · 2010
Earlier work this paper cites.
Nearly minimax optimal reward-free reinforcement learning
Zhang, Z · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Chu, W · 2011
Cited alongside, same era.
PAC bounds for discounted MDPs
Lattimore, T · 2012
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Cited alongside, same era.
Linear multi-resource allocation with semi-bandit feedback
Lattimore, T · 2015
Cited alongside, same era.
Model-based RL in contextual decision processes: PAC bounds and exponential improvements over model-free approaches
Sun, W · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Jia, Z · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C · 2020
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Li, G · 2020
Later among the works it cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A · 2020
Later among the works it cites.
Logarithmic regret for reinforcement learning with linear function approximation
He, J · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pac reinforcement learning with rich observations
Krishnamurthy, A · 2016
Cited alongside, same era.
Martingale inequalities of type dzhaparidze and van zanten
Fan, X · 2017
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
Jiang, N · 2017
Cited alongside, same era.
On oracle-efficient pac rl with rich observations
Dann, C · 2018
Cited alongside, same era.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Jiang, N · 2018
Cited alongside, same era.
Information directed sampling and bandits with heteroscedastic noise
Kirschner, J · 2018
Cited alongside, same era.
Later among the works it cites.
Improved regret analysis for variance-adaptive linear bandits and horizon-free linear mixture mdps
Kim, Y · 2021
Later among the works it cites.
Settling the horizon-dependence of sample complexity in reinforcement learning
Li, Y · 2021
Later among the works it cites.
Nearly horizon-free offline reinforcement learning
Ren, T · 2021
Later among the works it cites.
Stochastic shortest path: Minimax, parameter-free and towards horizon-free regret
Tarbouriech, J · 2021
Later among the works it cites.
Nearly optimal algorithms for linear contextual bandits with adversarial corruptions
He, J · 2022
Closest in time.
Horizon-free reinforcement learning in polynomial time: the power of stationary policies
Zhang, Z · 2022
Closest in time.