Fetching the paper…
Reading the bibliography…
We study reinforcement learning (RL) with linear function approximation where the underlying transition probability kernel of the Markov decision process (MDP) is a linear mixture model (Jia et al., 2020; Ayoub et al., 2020; Zhou et al., 2020) and the learning agent has access to either an integration or a sampling oracle of the individual basis kernels.
Zanette, A · 1901
Earlier work this paper cites.
Exploration-exploitation tradeoff using variance estimates in multi-armed bandits
Audibert, J.-Y · 1902
Earlier work this paper cites.
Tight regret bounds for infinite-armed linear contextual bandits
Li, Y · 1905
Earlier work this paper cites.
Near-optimal optimistic reinforcement learning using empirical Bernstein inequalities
Tossou, A · 1905
Earlier work this paper cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. F · 1905
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Cai, Q · 1912
Earlier work this paper cites.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y · 1912
Earlier work this paper cites.
Weighted sums of certain dependent random variables
Azuma, K · 1967
Earlier work this paper cites.
On tail probabilities for martingales
Freedman, D · 1975
Earlier work this paper cites.
Best linear unbiased estimation and prediction under a selection model
Henderson, C. R · 1975
Earlier work this paper cites.
Generalized polynomial approximations in Markovian decision processes
Schweitzer, P · 1985
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Earlier work this paper cites.
Improved optimistic algorithms for logistic bandits
Faury, L · 2002
Earlier work this paper cites.
Regret bounds for discounted mdps
Liu, S · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M · 2003
Earlier work this paper cites.
Learning near optimal policies with low inherent Bellman error
Zanette, A · 2003
Earlier work this paper cites.
Stochastic optimal control: the discrete-time case
Bertsekas, D. P · 2004
Earlier work this paper cites.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zhang, Z · 2004
Earlier work this paper cites.
Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?
Wang, R · 2005
Earlier work this paper cites.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2006
Earlier work this paper cites.
Q-learning with logarithmic regret
Yang, K · 2006
Cited alongside, same era.
Model-free reinforcement learning: from clipped pseudo-regret to sample complexity
Zhang, Z · 2006
Cited alongside, same era.
Provably efficient reinforcement learning for discounted MDPs with feature mapping
Zhou, D · 2006
Cited alongside, same era.
Stochastic linear optimization under bandit feedback
Dani, V · 2008
Cited alongside, same era.
Empirical Bernstein bounds and sample variance penalization
Maurer, A · 2009
Cited alongside, same era.
Policy error bounds for model-based reinforcement learning with factored linear models
Pires, B · 2016
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Azar, M. G · 2017
Later among the works it cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Jiang, N · 2017
Later among the works it cites.
On oracle-efficient pac rl with rich observations
Dann, C · 2018
Later among the works it cites.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Jiang, N · 2018
Later among the works it cites.
Is Q-learning provably efficient?
Jin, C · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, Z · 2009
Cited alongside, same era.
Logarithmic regret for reinforcement learning with linear function approximation
He, J · 2010
Cited alongside, same era.
Minimax optimal reinforcement learning for discounted MDPs
He, J · 2010
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, T · 2010
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Li, L · 2010
Cited alongside, same era.
Linearly parameterized bandits
Rusmevichientong, P · 2010
Cited alongside, same era.
Weisz, G · 2010
Cited alongside, same era.
Information directed sampling and bandits with heteroscedastic noise
Kirschner, J · 2018
Later among the works it cites.
Sidford, A · 2018
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Du, S. S · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular MDPs
Simchowitz, M · 2019
Later among the works it cites.
Model-based RL in contextual decision processes: PAC bounds and exponential improvements over model-free approaches
Sun, W · 2019
Later among the works it cites.
Regret minimization for reinforcement learning by evaluating the optimal bias function
Zhang, Z · 2019
Later among the works it cites.
Model-based reinforcement learning with a generative model is minimax optimal
Agarwal, A · 2020
Closest in time.
Model-based reinforcement learning with value-targeted regression
Jia, Z · 2020
Closest in time.
Provably efficient reinforcement learning with linear function approximation
Jin, C · 2020
Closest in time.
Bandit algorithms
Lattimore, T · 2020
Closest in time.
Learning with good feature representations in bandits and in rl with a generative model
Lattimore, T · 2020
Closest in time.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A · 2020
Closest in time.
A unifying view of optimism in episodic reinforcement learning
Neu, G · 2020
Closest in time.