Fetching the paper…
Reading the bibliography…
We study sample efficient reinforcement learning (RL) under the general framework of interactive decision making, which includes Markov decision process (MDP), partially observable Markov decision process (POMDP), and predictive state representation (PSR) as special cases.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y · 1912
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Weighted sums of certain dependent random variables
Azuma, K · 1967
Earlier work this paper cites.
The complexity of markov decision processes
Papadimitriou, C. H · 1987
Earlier work this paper cites.
Observable operator models for discrete stochastic time series
Jaeger, H · 2000
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Strens, M · 2000
Earlier work this paper cites.
Predictive representations of state
Littman, M · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S · 2002
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
From ε \varepsilon -entropy to kl-entropy: Analysis of minimum information complexity density estimation
Zhang, T · 2006
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Auer, P · 2008
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Rusmevichientong, P · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Earlier work this paper cites.
Closing the learning-planning loop with predictive state representations
Boots, B · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M · 2011
Earlier work this paper cites.
Predictive state representations: A new theory for modeling dynamical systems
Singh, S · 2012
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Russo, D · 2013
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
Osband, I · 2014
Earlier work this paper cites.
Learning to optimize via posterior sampling
Russo, D · 2014
Cited alongside, same era.
Probability in high dimension
Van Handel, R · 2014
Cited alongside, same era.
Supervised learning for dynamical system learning
Hefny, A · 2015
Cited alongside, same era.
Reinforcement learning of pomdps using spectral methods
Azizzadenesheli, K · 2016
Cited alongside, same era.
A pac rl algorithm for episodic pomdps
Guo, Z. D · 2016
Cited alongside, same era.
Pac reinforcement learning with rich observations
Krishnamurthy, A · 2016
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Provably efficient exploration in policy optimization
Cai, Q · 2020
Later among the works it cites.
Information theoretic regret bounds for online nonlinear control
Kakade, S · 2020
Later among the works it cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A · 2020
Later among the works it cites.
Deep dynamics models for learning dexterous manipulation
Nagabandi, A · 2020
Later among the works it cites.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Wang, R · 2020
Later among the works it cites.
A provably efficient model-free posterior sampling method for episodic reinforcement learning
Dann, C · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Agrawal, S · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C · 2017
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
Jiang, N · 2017
Cited alongside, same era.
Ensemble sampling
Lu, X · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K · 2018
Cited alongside, same era.
Bilinear classes: A structural framework for provable generalization in rl
Du, S · 2021
Later among the works it cites.
The statistical complexity of interactive decision making
Foster, D. J · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Jin, C · 2021
Later among the works it cites.
Rl for latent mdps: Regret guarantees and a lower bound
Kwon, J · 2021
Later among the works it cites.
Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning
Li, G · 2021
Later among the works it cites.
Ucb momentum q-learning: Correcting the bias without forgetting
Ménard, P · 2021
Later among the works it cites.
Sublinear regret for learning pomdps
Xiong, Y · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Zhou, D · 2021
Later among the works it cites.
Reinforcement learning from partial observation: Linear function approximation with provable sample efficiency
Cai, Q · 2022
Closest in time.
Provable reinforcement learning with a short-term memory
Efroni, Y · 2022
Closest in time.
Embed to control partially observed systems: Representation learning with provable sample efficiency
Wang, L · 2022
Closest in time.
Nearly optimal policy optimization with stable at any time guarantee
Wu, T · 2022
Closest in time.
A self-play posterior sampling algorithm for zero-sum Markov games
Xiong, W · 2022
Closest in time.
Pac reinforcement learning for predictive state representations
Zhan, W · 2022
Closest in time.
Horizon-free reinforcement learning in polynomial time: the power of stationary policies
Zhang, Z · 2022
Closest in time.