Fetching the paper…
Reading the bibliography…
We consider the problem of learning in episodic finite-horizon Markov decision processes with an unknown transition function, bandit feedback, and adversarial losses.
Optimal adaptive policies for markov decision processes
Burnetas, A. N. and Katehakis, M. N · 1997
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
Hannan consistency in on-line learning in case of unbounded losses under partial monitoring
Allenberg, C., Auer, P., Györfi, L., and Ottucsák, G · 2006
Earlier work this paper cites.
Competing in the dark: An efficient algorithm for bandit linear optimization
Abernethy, J. D., Hazan, E., and Rakhlin, A · 2008
Earlier work this paper cites.
Online markov decision processes
Even-Dar, E., Kakade, S. M., and Mansour, Y · 2009
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Maurer, A. and Pontil, M · 2009
Earlier work this paper cites.
Arbitrarily modulated markov decision processes
Yu, J. Y. and Mannor, S · 2009
Earlier work this paper cites.
Markov decision processes with arbitrary reward processes
Yu, J. Y., Mannor, S., and Shimkin, N · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
The online loop-free stochastic shortest-path problem
Neu, G., György, A., and Szepesvári, C · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Cited alongside, same era.
Contextual bandit algorithms with supervised learning guarantees
Beygelzimer, A., Langford, J., Li, L., Reyzin, L., and Schapire, R · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Chu, W., Li, L., Reyzin, L., and Schapire, R · 2011
Cited alongside, same era.
Deterministic mdps with adversarial rewards and bandit feedback
Arora, R., Dekel, O., and Tewari, A · 2012
Cited alongside, same era.
The adversarial stochastic shortest path problem with unknown transition probabilities
Neu, G., Gyorgy, A., and Szepesvari, C · 2012
Cited alongside, same era.
Better rates for any adversarial deterministic mdp
Dekel, O. and Hazan, E · 2013
Cited alongside, same era.
Introduction to online convex optimization
Hazan, E. et al · 2016
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Later among the works it cites.
Learning unknown markov decision processes: a thompson sampling approach
Ouyang, Y., Gagrani, M., Nayyar, A., and Jain, R · 2017
Later among the works it cites.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Fruit, R., Pirotta, M., Lazaric, A., and Ortner, R · 2018
Later among the works it cites.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Later among the works it cites.
Reinforcement learning under drift
Cheung, W. C., Simchi-Levi, D., and Zhu, R · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Online learning in episodic markovian decision processes by relative entropy policy search
Zimin, A. and Neu, G · 2013
Cited alongside, same era.
Bandits with switching costs: T 2 / 3 {T}^{2/3} regret
Dekel, O., Ding, J., Koren, T., and Peres, Y · 2014
Cited alongside, same era.
Online markov decision processes under bandit feedback
Neu, G., Antos, A., György, A., and Szepesvári, C · 2014
Cited alongside, same era.
Explore no more: Improved high-probability regret bounds for non-stochastic bandits
Neu, G · 2015
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P
Cited in the paper.
The nonstochastic multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E
Cited in the paper.
Closest in time.
Corruption robust exploration in episodic reinforcement learning
Lykouris, T., Simchowitz, M., Slivkins, A., and Sun, W · 2019
Closest in time.
Q-learning with ucb exploration is sample efficient for infinite-horizon mdp
Wang, Y., Dong, K., Chen, X., and Wang, L · 2019
Closest in time.
Model-free reinforcement learning in infinite-horizon average-reward markov decision processes
Wei, C.-Y., Jafarnia-Jahromi, M., Luo, H., Sharma, H., and Jain, R · 2019
Closest in time.
Regret minimization for reinforcement learning by evaluating the optimal bias function
Zhang, Z. and Ji, X · 2019
Closest in time.