Fetching the paper…
Reading the bibliography…
Learning Markov decision processes (MDPs) in the presence of the adversary is a challenging problem in reinforcement learning (RL).
Abbasi-Yadkori, Y · 1908
Earlier work this paper cites.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y · 1912
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L · 1994
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S · 1999
Earlier work this paper cites.
Direct gradient-based reinforcement learning
Baxter, J · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2001
Earlier work this paper cites.
Optimistic policy optimization with bandit feedback
Efroni, Y · 2002
Earlier work this paper cites.
Provably efficient adaptive approximate policy iteration
Hao, B · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S · 2002
Earlier work this paper cites.
Covariant policy search
Bagnell, J. A · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M · 2003
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N · 2006
Earlier work this paper cites.
Online markov decision processes
Even-Dar, E · 2009
Cited alongside, same era.
Markov decision processes with arbitrary reward processes
Yu, J. Y · 2009
Cited alongside, same era.
Online markov decision processes under bandit feedback
Gergely Neu, A. G · 2010
Cited alongside, same era.
The online loop-free stochastic shortest-path problem
Neu, G · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Cited alongside, same era.
The adversarial stochastic shortest path problem with unknown transition probabilities
Neu, G · 2012
Cited alongside, same era.
Better rates for any adversarial deterministic mdp
Is q-learning provably efficient?
Jin, C · 2018
Later among the works it cites.
Information directed sampling and bandits with heteroscedastic noise
Kirschner, J · 2018
Later among the works it cites.
Provably efficient rl with rich observations via latent state decoding
Du, S · 2019
Later among the works it cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A · 2019
Later among the works it cites.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Sun, W · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dekel, O · 2013
Cited alongside, same era.
Online learning in episodic markovian decision processes by relative entropy policy search
Zimin, A · 2013
Cited alongside, same era.
Scaling up robust mdps using function approximation
Tamar, A · 2014
Cited alongside, same era.
Trust region policy optimization
Schulman, J · 2015
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N · 2017
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Neu, G · 2017
Cited alongside, same era.
Provably efficient exploration in policy optimization
Cai, Q · 2020
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Jia, Z · 2020
Later among the works it cites.
On the stability and convergence of robust adversarial reinforcement learning: A case study on linear quadratic systems
Zhang, K · 2020
Later among the works it cites.
Logarithmic regret for reinforcement learning with linear function approximation
He, J · 2021
Closest in time.
Online robust reinforcement learning with model uncertainty
Wang, Y · 2021
Closest in time.
Nearly minimax optimal regret for learning infinite-horizon average-reward mdps with linear function approximation
Wu, Y · 2022
Closest in time.