Fetching the paper…
Reading the bibliography…
We study episodic reinforcement learning under unknown adversarial corruptions in both the rewards and the transition probabilities of the underlying system.
The nonstochastic multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E · 2002
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M · 2002
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Earlier work this paper cites.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Earlier work this paper cites.
Stochastic bandits robust to adversarial corruptions
Lykouris, T., Mirrokni, V., and Paes Leme, R · 2018
Earlier work this paper cites.
Exploration in structured reinforcement learning
Ok, J., Proutiere, A., and Tranos, D · 2018
Earlier work this paper cites.
Better algorithms for stochastic bandits with adversarial corruptions
Gupta, A., Koren, T., and Talwar, K · 2019
Earlier work this paper cites.
Stochastic linear optimization with adversarial corruption
Li, Y., Lou, E. Y., and Shan, L · 2019
Cited alongside, same era.
Data poisoning attacks on stochastic bandits
Liu, F. and Shroff, N · 2019
Cited alongside, same era.
Online convex optimization in adversarial markov decision processes
Rosenberg, A. and Mansour, Y · 2019
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M. and Jamieson, K. G · 2019
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Zanette, A. and Brunskill, E · 2019
Cited alongside, same era.
An optimal algorithm for stochastic and adversarial bandits
Simultaneously learning stochastic and adversarial episodic mdps with known transition
Jin, T. and Luo, H · 2020
Later among the works it cites.
Adaptive reward-free exploration, 2020
Kaufmann, E., Ménard, P., Domingues, O. D., Jonsson, A., Leurent, E., and Valko, M · 2020
Later among the works it cites.
Bias no more: high-probability data-dependent regret bounds for adversarial bandits and mdps
Lee, C.-W., Luo, H., Wei, C.-Y., and Zhang, M · 2020
Later among the works it cites.
Corruption robust exploration in episodic reinforcement learning, 2020
Lykouris, T., Simchowitz, M., Slivkins, A., and Sun, W · 2020
Later among the works it cites.
Fast active learning for pure exploration in reinforcement learning, 2020
Ménard, P., Domingues, O. D., Jonsson, A., Kaufmann, E., Leurent, E., and Valko, M · 2020
Later among the works it cites.
Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zimmert, J. and Seldin, Y · 2019
Cited alongside, same era.
Stochastic linear bandits robust to adversarial attacks, 2020
Bogunovic, I., Losalka, A., Krause, A., and Scarlett, J · 2020
Cited alongside, same era.
Learning adversarial Markov decision processes with bandit feedback and unknown transition
Jin, C., Jin, T., Luo, H., Sra, S., and Yu, T · 2020
Cited alongside, same era.
Wang, R., Du, S. S., Yang, L. F., and Kakade, S. M · 2020
Later among the works it cites.
Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon, 2020
Zhang, Z., Ji, X., and Du, S. S · 2020
Later among the works it cites.
Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap
Xu, H., Ma, T., and Du, S. S · 2021
Closest in time.