Fetching the paper…
Reading the bibliography…
We study episodic reinforcement learning (RL) in non-stationary linear kernel Markov decision processes (MDPs).
Sample-optimal parametric q-learning using linearly additive features
Yang, L. F · 1902
Earlier work this paper cites.
Hedging the drift: Learning to optimize under non-stationarity
Cheung, W. C · 1903
Earlier work this paper cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. F · 1905
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C · 1907
Earlier work this paper cites.
On the global convergence of actor-critic: A case for linear quadratic regulator with ergodic cost
Yang, Z · 1907
Earlier work this paper cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A · 1908
Earlier work this paper cites.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L · 1909
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
Akkaya, I · 1910
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Cai, Q · 1912
Earlier work this paper cites.
Learning adversarial mdps with bandit feedback and unknown transition
Jin, C · 1912
Earlier work this paper cites.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y · 1912
Earlier work this paper cites.
Probabilistic computations: Toward a unified measure of complexity
Yao, A. C.-C · 1977
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovsky, A. S · 1983
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Complexity analysis of real-time reinforcement learning
Koenig, S · 1993
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J · 1996
Earlier work this paper cites.
Optimistic policy optimization with bandit feedback
Efroni, Y · 2002
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2002
Earlier work this paper cites.
Provably efficient reinforcement learning with general value function approximation
Wang, R · 2005
Earlier work this paper cites.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2006
Earlier work this paper cites.
Reinforcement learning for non-stationary markov decision processes: The blessing of (more) optimism
Cheung, W. C · 2006
Earlier work this paper cites.
Provably efficient reinforcement learning for discounted MDPs with feature mapping
Zhou, D · 2006
Earlier work this paper cites.
PC-PG: Policy cover directed exploration for provable policy gradient learning
Agarwal, A · 2007
Earlier work this paper cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Cen, S · 2007
Earlier work this paper cites.
A kernel-based approach to non-stationary reinforcement learning in metric spaces
Domingues, O. D · 2007
Earlier work this paper cites.
Dynamic regret of policy optimization in non-stationary environments
Fei, Y · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Auer, P · 2009
Earlier work this paper cites.
Online markov decision processes
Even-Dar, E · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T · 2010
Cited alongside, same era.
The online loop-free stochastic shortest-path problem
Neu, G · 2010
Cited alongside, same era.
Linearly parameterized bandits
Rusmevichientong, P · 2010
Cited alongside, same era.
Efficient learning in non-stationary linear markov decision processes
Touati, A · 2010
Cited alongside, same era.
Nonstationary reinforcement learning with linear function approximation
Zhou, H · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Cited alongside, same era.
Robust deep reinforcement learning with adversarial attacks
Pattanaik, A · 2017
Later among the works it cites.
Robust adversarial reinforcement learning
Pinto, L · 2017
Later among the works it cites.
Deep reinforcement learning framework for autonomous driving
Sallab, A. E · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On upper-confidence bound policies for switching bandit problems
Garivier, A · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S · 2012
Cited alongside, same era.
The best of both worlds: Stochastic and adversarial bandits
Bubeck, S · 2012
Cited alongside, same era.
Pac bounds for discounted mdps
Lattimore, T · 2012
Cited alongside, same era.
The adversarial stochastic shortest path problem with unknown transition probabilities
Neu, G · 2012
Cited alongside, same era.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Zhou, D · 2012
Cited alongside, same era.
Gajane, P · 2018
Later among the works it cites.
Is q-learning provably efficient?
Jin, C · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Later among the works it cites.
More adaptive algorithms for adversarial bandits
Wei, C.-Y · 2018
Later among the works it cites.
Optimal exploration–exploitation in a multi-armed bandit problem with non-stationary rewards
Besbes, O · 2019
Later among the works it cites.
A new algorithm for non-stationary contextual bandits: Efficient, optimal and parameter-free
Chen, Y · 2019
Later among the works it cites.
Neural trust region/proximal policy optimization attains globally optimal policy
Liu, B · 2019
Later among the works it cites.
Online convex optimization in adversarial markov decision processes
Rosenberg, A · 2019
Later among the works it cites.
Weighted linear bandits for non-stationary environments
Russac, Y · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, O. M · 2020
Later among the works it cites.
Simultaneously learning stochastic and adversarial episodic mdps with known transition
Jin, T · 2020
Later among the works it cites.
Bandit algorithms
Lattimore, T · 2020
Later among the works it cites.
Near-optimal regret bounds for model-free rl in non-stationary episodic mdps
Mao, W · 2020
Later among the works it cites.
On the global convergence rates of softmax policy gradient methods
Mei, J · 2020
Later among the works it cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A · 2020
Later among the works it cites.
Variational regret bounds for reinforcement learning
Ortner, R · 2020
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Shani, L · 2020
Later among the works it cites.
Learning near optimal policies with low inherent bellman error
Zanette, A · 2020
Later among the works it cites.
A simple approach for non-stationary linear bandits
Zhao, P · 2020
Later among the works it cites.
The best of both worlds: stochastic and adversarial episodic mdps with unknown transition
Jin, T · 2021
Closest in time.
Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach
Wei, C.-Y · 2021
Closest in time.
Non-stationary linear bandits revisited
Zhao, P · 2021
Closest in time.