Fetching the paper…
Reading the bibliography…
We study reinforcement learning (RL) with linear function approximation under the adaptivity constraint.
Sequential batch learning in finite-action linear contextual bandits
Han, Y · 2004
Earlier work this paper cites.
Efficient algorithms for online decision problems
Kalai, A · 2005
Earlier work this paper cites.
Linear bandits with limited adaptivity and learning distributional optimal design
Ruan, Y · 2007
Earlier work this paper cites.
Regret minimization for online buffering problems using the weighted majority algorithm
Geulen, S · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Earlier work this paper cites.
Online bandit learning against an adaptive adversary: from regret to policy regret
Arora, R · 2012
Earlier work this paper cites.
Online learning with switching costs and other adaptive adversaries
Cesa-Bianchi, N · 2013
Earlier work this paper cites.
Bandits with switching costs: T 2/3 regret
Dekel, O · 2014
Earlier work this paper cites.
Random-walk perturbations for online combinatorial optimization
Devroye, L · 2015
Earlier work this paper cites.
Batched bandit problems
Perchet, V · 2016
Earlier work this paper cites.
Online learning over a finite action set with limited switching
Altschuler, J · 2018
Earlier work this paper cites.
Provably efficient q-learning with low switching cost
Bai, Y · 2019
Cited alongside, same era.
Batched multi-armed bandits problem
Gao, Z · 2019
Cited alongside, same era.
Consistent online optimization: Convex and submodular
Jaghargh, M. R. K · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Yang, L · 2019
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Cai, Q · 2020
Cited alongside, same era.
Minimax regret of switching-constrained online convex optimization: No phase transition
A unifying view of optimism in episodic reinforcement learning
Neu, G · 2020
Later among the works it cites.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Wang, R · 2020
Later among the works it cites.
On function approximation in reinforcement learning: Optimism in the face of large state spaces
Yang, Z · 2020
Later among the works it cites.
Learning near optimal policies with low inherent bellman error
Zanette, A · 2020
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in rl
Du, S. S · 2021
Closest in time.
Regret bounds for batched bandits
Esfandiari, H · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, L · 2020
Cited alongside, same era.
Is a good representation sufficient for sample efficient reinforcement learning?
Du, S. S · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Jia, Z · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C · 2020
Cited alongside, same era.
Learning with good feature representations in bandits and in rl with a generative model
Lattimore, T · 2020
Cited alongside, same era.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Zhou, D
Cited in the paper.
Instance-dependent complexity of contextual bandits and reinforcement learning: A disagreement-based perspective
Foster, D. J · 2021
Closest in time.
A provably efficient algorithm for linear markov decision process with low switching cost
Gao, M · 2021
Closest in time.
Logarithmic regret for reinforcement learning with linear function approximation
He, J · 2021
Closest in time.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y · 2021
Closest in time.
Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon
Zhang, Z · 2021
Closest in time.